Matthew Rosenzweig, Dejan Slepčev, Lihan Wangmath.AP math.PR stat.ML
We study the Wasserstein gradient flow of the squared Maximum Mean Discrepancy (MMD) generated by the nonsmooth energy kernels $K(z)=-|z|^q$, $0<q<2$. In dimensions $d\ge2$, the corresponding energies are not displacement semiconvex, so standard Wasserstein-gradient-flow theory does not apply. When $d+q-2>0$, we prove global well-posedness on $\mathbb{R}^d$ for probability densities in subcritical $L^p$ spaces, with targets in the same integrability class and with finite moments. We also include the one-dimensional Coulomb endpoint $d=q=1$. For the associated $N$-particle system, we prove global noncollision and fixed-$N$ convergence to the collision-free critical set, a particle-to-continuum criticality principle, and a modulated-energy mean-field estimate that yields convergence of the particle dynamics to the continuum flow as $N\to\infty$ on every finite time interval. We also construct collision-free saddle equilibria, showing that deterministic particle trajectories need not approach global empirical minimizers. For $1\le q<2$, every continuum solution in our class has a narrowly relatively compact orbit, every $ω$-limit point is Lagrangian critical, and the orbit approaches the Lagrangian critical set. For $0<q<1$, the same conclusions hold under uniform-in-time moment and subcritical $L^p$ bounds. We prove that an absolutely continuous Lagrangian critical point equals the target when the source and target have finite moments of order $q$, except when $0<q<1$ and $d\in\{1,3\}$. Under the preceding uniform bounds, rigidity gives convergence of the continuum flow to the target throughout the rigid part of the well-posedness range. Finally, we show that no initial-data-independent multiplicative MMD decay modulus exists on $\mathbb{R}^d$, and that global Polyak--Łojasiewicz inequalities fail in several whole-space and periodic Riesz/Coulomb regimes.
Antonin Chodron de Courcel, Matthew Rosenzweigmath.AP stat.ML
We study the long-time behavior of the Wasserstein gradient flow of the squared Maximum Mean Discrepancy (MMD) between a probability measure $ρ$ and a target measure $μ$, where the underlying kernel is given by a Coulomb potential. For $L^\infty$ target densities $μ$, we establish the existence of global weak solutions starting from arbitrary Borel probability measures and prove that the density $ρ_t$ belongs to $L^\infty$ for any $t>0$. We also show that the Hölder norm can grow exponentially in time. On the flat torus ${\mathbb{T}}^\mathsf{d}$, we prove a global metric PL inequality for every finite-Coulomb-energy source and nearly uniform target. For general bounded, uniformly positive targets, we prove exponential decay of the squared MMD without requiring a lower bound on the initial data, using a defective PL inequality. We also prove that the usual PL inequality may fail when the target vanishes only at one point and that, when $\mathsf{d}\ge2$, no PL constant can hold uniformly over all targets satisfying a prescribed lower bound. On ${\mathbb{R}}^\mathsf{d}$, for $\mathsf{d}\ge2$, under radial symmetry, source-support inclusion, and target-positivity assumptions, we establish a PL inequality and exponential convergence. On the unrestricted whole-space class, neither a multiplicative squared-MMD decay modulus uniform over the initial datum nor a global PL inequality can hold. Finally, in every dimension and in both spatial settings, we prove that every Lagrangian critical point coincides with the target when $(ρ-μ)^+$ is absolutely continuous. In dimension two, the energy supplies uniform tightness. This implies that if our constructed solutions have finite energy at some positive time, then they converge to the target narrowly and strongly in negative-order Sobolev spaces.
Markus Heinonen, Yair Shenfeld, Ricardo Baptista +4stat.ML cs.AI cs.LG
Reconstructing population dynamics is a central problem in the physical and data sciences. Often, the dynamics are modeled as a Wasserstein gradient flow (WGF): a curve of distributions driven by an energy functional. Though there are multiple mathematical characterizations of a WGF, the dominant algorithmic approach relies on the Jordan--Kinderlehrer--Otto (JKO) scheme. JKO-based methods are inflexible to time discretisation and require solving costly optimal transport problems. We take a residual approach, enforcing the continuity equations via a non-negative loss function whose minimum is the WGF. Combined with a data-fitting divergence, this gives a single global objective. This perspective unifies several existing methods and leads to a new particle-based method, stitching, that is simulation-free and robust to large gaps between observations. We demonstrate that the stitching method achieves state-of-the-art performance across trajectory inference benchmarks. For code see github.com/BasisResearch/wasserstein-residuals.
Sophia Seulkee Kang, Louis Sharrock, Xiaoyuan Cheng +2cs.LG stat.ML
Minimum maximum mean discrepancy (MMD) estimation has emerged as a robust and likelihood-free alternative to maximum likelihood estimation for parameter estimation. Yet, despite its practical success, the associated optimization problem remains poorly understood, with theoretical guarantees for existing algorithms hinging on convexity assumptions that rarely hold in practice. We address this gap by proposing a preconditioned gradient descent (PGD) scheme, establishing its asymptotic \emph{global} convergence under explicit gradient-dominance and projection-residual conditions. Our approach is inspired by recent progress on MMD gradient flows, a nonparametric descent scheme on the space of probability measures. We provide extensive empirical evidence that our PGD scheme outperforms standard gradient descent across a range of challenging parameter estimation and composite hypothesis testing problems.