Matthew Rosenzweig, Dejan Slepčev, Lihan Wangmath.AP math.PR stat.ML
We study the Wasserstein gradient flow of the squared Maximum Mean Discrepancy (MMD) generated by the nonsmooth energy kernels $K(z)=-|z|^q$, $0<q<2$. In dimensions $d\ge2$, the corresponding energies are not displacement semiconvex, so standard Wasserstein-gradient-flow theory does not apply. When $d+q-2>0$, we prove global well-posedness on $\mathbb{R}^d$ for probability densities in subcritical $L^p$ spaces, with targets in the same integrability class and with finite moments. We also include the one-dimensional Coulomb endpoint $d=q=1$. For the associated $N$-particle system, we prove global noncollision and fixed-$N$ convergence to the collision-free critical set, a particle-to-continuum criticality principle, and a modulated-energy mean-field estimate that yields convergence of the particle dynamics to the continuum flow as $N\to\infty$ on every finite time interval. We also construct collision-free saddle equilibria, showing that deterministic particle trajectories need not approach global empirical minimizers. For $1\le q<2$, every continuum solution in our class has a narrowly relatively compact orbit, every $ω$-limit point is Lagrangian critical, and the orbit approaches the Lagrangian critical set. For $0<q<1$, the same conclusions hold under uniform-in-time moment and subcritical $L^p$ bounds. We prove that an absolutely continuous Lagrangian critical point equals the target when the source and target have finite moments of order $q$, except when $0<q<1$ and $d\in\{1,3\}$. Under the preceding uniform bounds, rigidity gives convergence of the continuum flow to the target throughout the rigid part of the well-posedness range. Finally, we show that no initial-data-independent multiplicative MMD decay modulus exists on $\mathbb{R}^d$, and that global Polyak--Łojasiewicz inequalities fail in several whole-space and periodic Riesz/Coulomb regimes.
Infinite-width limits are a standard way to reason about neural networks, but it is not automatic that the limiting learner has the same complexity-theoretic inductive bias as large finite networks. We study this question for Bayesian neural networks at the mean-field, or critical feature-learning, scaling. The central quantity is the \emph{reduced entropy} \[ s_\infty(y,\varepsilon)=\limsup_N -\frac{1}{N}\log π_N^0(L\le \varepsilon), \] the intensive prior cost of representing a target function $y$ to population mean-squared error $\varepsilon$. Our main result is a width-robust learnability theorem. At fixed depth, a family of Boolean-cube targets is learnable from polynomially many samples at infinite width if and only if it is learnable at polynomial width, if and only if its reduced entropy is polynomially bounded. Equivalently, up to polynomial slack in accuracy, the Bayesian mean-field learner generalizes exactly on the targets that can be represented by polynomial-size networks. The forward direction is proved by a form of subsampling: from the infinitely many hidden neurons in the mean-field solution, one can select polynomially many representatives and still preserve the learned function on every input simultaneously. At the critical scaling this subsampling has both an ``active'' component, which keeps the data-dependent low-dimensional statistics, and a ``lazy'' component, which resamples the entropy-dominated directions from the prior. Thus the infinite-width mean-field limit gives a clean analytic description of learning without introducing spurious width-dependent generalization power.
Antonin Chodron de Courcelmath.OC cs.AI cs.LG math-ph math.AP
We study the dynamics of gradient descent in the Edge of Stability regime, where the learning rate is large enough to induce persistent oscillations in the loss and the sharpness. We propose a continuous-time effective model that tracks the evolution of the average trajectory coupled with the time-averaged covariance of its fast oscillations. Our analysis reveals that the natural quantity to monitor in such unstable regimes is an effective free energy, which combines the original risk functional with a curvature-related "entropic" term. Our model allows us to track the envelope of the oscillations even in situations where its dynamics evolve on similar timescales as the averaged weights. Otherwise stated, we can track the spikes that occur during the training of some neural network architectures. For wide two-layer neural networks optimized under stable non-vanishing oscillations, we derive a mean-field limit that results in a novel kinetic equation describing the joint distribution of weights and their fluctuations. We show that this equation can be interpreted as a Wasserstein-2 gradient flow of a macroscopic free energy. Finally, we provide numerical evidence on matrix factorization and deep learning tasks (CIFAR-10) to demonstrate the model's accuracy in capturing the envelope of the oscillations and the predictive power of the effective free energy.