We present the mathematical foundations of linear independent component analysis (ICA) models based on standard literature in a self-contained note. It is aimed at readers with a background in measure-theoretic probability theory. We first develop the theory of the characteristic functions of probability measures on $\mathbb{R}^d$, including their analyticity and the way in which they determine and characterise the distributions. We then focus on several identifiability results of ICA models with successively strengthened assumptions on the sources: from merely non-constant, to non-Gaussian, to Gaussian-free independent sources. Under the strictest assumptions, we show that the independent sources are identifiable up to translation, permutation, scales and signs, and this even in the presence of additive Gaussian noise. Furthermore, we present the online equivariant gradient descent ICA algorithm for recovering the independent sources from data, in the standard complete noiseless non-Gaussian ICA setting.
The Mahalanobis distance is a fundamental covariance-adapted metric for multivariate data and plays a central role in recovering latent geometry from nonlinear observations. We extend this principle from vector-valued data to probability measures by introducing a Wasserstein Mahalanobis distance. Our construction replaces Euclidean displacement vectors with optimal transport displacement fields and local covariance matrices with covariance operators defined on Wasserstein tangent spaces. We show that this construction inherits the geometry-recovery property underlying nonlinear independent component analysis. In particular, for Gaussian measures with common covariance transformed by a smooth nonlinear pushforward, the proposed Wasserstein Mahalanobis distance approximates the classical Mahalanobis distance between the transformed latent means. The correspondence is exact for affine transformations and holds up to controlled higher-order error terms for general smooth transformations. These results establish a distribution-valued analog of classical Mahalanobis geometry and provide theoretical support for covariance-adapted learning directly in Wasserstein space. Numerical experiments confirm the theoretical predictions and demonstrate accurate recovery of latent geometric structure.
Ashutosh Jha, Michel Besserve, Simon Buchholzcs.LG stat.ML
Linear Independent Component Analysis (ICA) recovers jointly independent source signals from their linear mixtures. To achieve this, classical ICA algorithms attempt to maximize non-Gaussianity, measured by negentropy, which is linked to independence by information theory. Because exact negentropy optimization is intractable, they rely on proxy contrast functions, such as fourth-order cumulants, and parametric log-likelihoods. We propose instead to measure non-Gaussianity using the squared Wasserstein distance $W_2^2$ to a standard Gaussian. We prove that the Wasserstein distance between a standard normal distribution and linear projections of the data is maximized when the projection recovers an independent component. Based on this observation, we propose the OT-ICA algorithm which finds this projection by gradient-based optimization. Empirical evaluation on simulated data shows that OT-ICA outperforms proxy-based methods for different distributions of the latent variables. Application to EEG artifact removal and econometric price discovery confirm OT-ICA can be used for applied ICA tasks without distributional assumptions.
Félix Laplante, Christophe Ambroise, Pierre Humbertstat.ML stat.ME
We study the squared $2$-Wasserstein distance to the standard Gaussian as a non-Gaussianity criterion and use it for linear Independent Component Analysis (ICA) and causal inference in Linear Non-Gaussian Acyclic Models (LiNGAM). The analysis relies on a strict inequality between the Wasserstein non-Gaussianity of independent standardized sources and that of their linear combinations. When at most one source is Gaussian, any unit-norm linear combination involving at least two sources has strictly smaller squared Wasserstein distance than the corresponding weighted sum of source distances. At the population level, this yields exact identification of the ICA unmixing matrix, up to signed permutation, and gives an analogous characterization of causal orders through least-squares residuals. We then define empirical plug-in estimators and prove distribution-free uniform convergence bounds under finite-moment assumptions, before detailing three practical solvers: a Picard-style orthogonal optimizer for ICA, an exhaustive dynamic program for causal-order search, and a greedy order-search variant. Empirically, we demonstrate competitive performance for both tasks and provide open-source implementations for source separation and causal inference.
There is a gap between the theoretical foundations of disentanglement and the practice of modern representation learning. Existing theoretical frameworks, particularly Independent Component Analysis (ICA) and its nonlinear variants, assume a generative model with statistically independent latent variables underlying the data so that disentanglement amounts to identifying the latents that could have generated the data. This generative framework is interpretable and theoretically justified, but its strong assumptions make it difficult to apply to modern representation learning. Modern pretrained encoders often learn features that exhibit disentangled properties without making generative assumptions, yet there is no general theory for interpreting these features as independent factors of variation. We take a step toward such a theory by introducing Riemannian ICA (RICA), which replaces ICA's global generative model with local geometric structure. RICA is founded on the observation that in ICA, the factors of variation underlying a data point can be understood through radial curves emanating from the point that map to axis-aligned lines in the latent space. We formalize this perspective using Riemannian geometry and introduce our theory in a way that is consistent with the existing generative approach. Our main contribution is the disentanglement tensor, which encodes a second-order notion of disentanglement that we call pointwise disentanglement. This tensor depends on the Hessian of the data log likelihood as well as the Ricci curvature induced by the model. In a controlled source recovery setting with known ground-truth sources, RICA recovers sources across several manifolds, while the success of ICA baselines depends on the coordinates used to represent the observations. Our work provides a theoretical basis for studying local disentanglement without assuming a global generative model.