Sagnik Nandy, Samriddha Lahiry, Pragya Sur +1stat.ME cs.LG math.ST stat.ML
Multimodal supervised learning seeks to leverage multiple heterogeneous data sources to improve predictive performance. A central challenge is determining the fusion granularity across modalities: over-integration may amplify noise while under-integration fails to exploit cross-modal dependence. Existing approaches rely on pre-specified fusion architectures, from early to late fusion, that may not adapt to the underlying dependence structure among modalities. We propose DAIF, a data adaptive intermediate fusion framework that combines random matrix theory and non-parametric dependence measures to learn fusion structure directly from data. We operate under a Bayesian multimodal factor model where the prior on the latent factors determines the cross-modal dependence. Our method clusters modalities based on estimated intermodal dependence, then performs clusterwise empirical Bayes estimation of the priors. These estimated priors are used to construct denoisers within an approximate message passing (AMP) framework, yielding denoised low-dimensional features that borrow strength across related modalities while preserving modality-specific signal. The resulting embeddings are used for downstream supervised prediction. We evaluate the framework through simulations under varying dependence structures and signal regimes, comparing against several benchmark methods, and demonstrate its practical utility on two multimodal datasets, namely a trimodal TEA-seq dataset (Swanson et al., 2021) and TCGA-BRCA dataset (Goldman et al., 2020). In the first example, we predict the expression level of a T-cell differentiation marker protein and in the second case we analyze patient survival prediction based on multimodal information. Our method competes with or outperforms the state-of-the-art techniques in both prediction problems, demonstrating its versatility across diverse supervised learning tasks.
Kevin Han Huang, Haoyu Ye, Somak Laha +1math.ST cs.LG stat.ML
Over-parameterized linear regression has been widely studied over the last decade. However, most existing works assume that the covariates are independent and that their covariance matrices are non-degenerate. In this paper, we relax both assumptions and derive deterministic equivalents for the prediction risk in a vanishing-ridge regime. We show that degeneracy of the covariance matrices and dependence can lead to multiple descent, and characterize where the corresponding peaks can occur. Our proofs use a novel graph representation of the variance profile. We show that maximum matchings and the Dulmage--Mendelsohn decomposition of the associated bipartite graph identify the configurations at which the variance becomes singular.
Principal component regression (PCR) regularizes high-dimensional prediction by choosing a spectral cutoff, but rank selection cannot correct systematic inflation of the retained empirical eigenvalues. We study clean Gaussian random designs in which the aggregate covariance tail creates a nearly scalar sample-space floor comparable to the predictive head scale. De-floored principal component regression (dPCR) retains the cutoff and subtracts an estimated floor from the retained denominators. We prove an ordinary-PCR prediction-risk lower bound uniform over all ranks and a high-probability dPCR upper bound. When the floor is sharp and inexpensive to remove in population prediction risk, the conditional risk of dPCR is asymptotically negligible relative to that of the best ordinary PCR rank. An exact risk decomposition explains the separation: denominator inflation is governed by first spectral mass, whereas the clean prediction cost of correction is governed by squared spectral mass. A same-sample trimmed-mean floor estimate attains the oracle dPCR upper-bound rate at a prespecified rank, and the separation persists under approximate predictive alignment when the tail prediction-energy fraction vanishes. Separate pointwise fixed-aspect formulas show that the risk-optimal positive scalar correction improves rank-$1$ PCR, whereas mean-floor subtraction is generally not optimal for a broad Marchenko--Pastur bulk.
We study the sample covariance error of centered Gaussians. A remarkable breakthrough [66] established the correct error scaling order and explicitly revealed the critical role of both the effective rank and the true covariance spectrum. In this work, we move beyond scaling characterizations and determine the precise limiting value of the error's spectral norm. To do so, we develop a generic framework based on Random Duality Theory (RDT). Within this framework, we first determine closed-form, explicit RDT-based upper bounds. We then establish complementary lower bounds by introducing a novel bilinear-quadratic RDT lower-bounding mechanism. By combining this mechanism with a two-replica systems bounding strategy, we show that our lower and upper bounds match in large-dimensional contexts. Our theoretical results are supplemented with numerical evaluations and simulations, demonstrating an excellent agreement already for problem sizes on the order of thousands.
Restricted Boltzmann machines (RBMs) represent data by shaping an energy landscape over visible and hidden configurations, but their discriminative use is fragile under out-of-distribution (OOD) inputs: samples outside the training distribution can be absorbed into one of the learned class basins rather than rejected. Here, we analyze this failure mode through the spectrum of the induced visible--visible interaction $J=WW^{T}$, where \(W\) is the visible--hidden weight matrix. Relative to a Marchenko--Pastur random-matrix reference, conventional training spreads spectral weight into many weak, bulk-compatible directions, increasing the effective rank of $J$. When auxiliary random binary images are assigned to a rejection label during training, the learned interaction undergoes effective-rank collapse: weak bulk-like modes are depleted, spectral weight concentrates into fewer dominant eigendirections, and the effective rank of $J$ approaches that of the empirical data covariance matrix. The resulting RBM rejects structured OOD image datasets while preserving MNIST classification accuracy, showing that random auxiliary exposure can reshape both the interaction spectrum and the free-energy landscape of an energy-based classifier.
Changwon Yoon, Minwoo Kim, Sungkyu Jung +1stat.ME stat.AP stat.ML
Statistical analysis of high-dimensional data is often hampered by limited sample sizes, yet auxiliary datasets from related sources are often readily available. When two such datasets share part of their covariance structure, but not all of it, exploiting the shared part can substantially improve estimation. We propose a spiked covariance model that explicitly captures this partial sharing: two datasets share a subspace of unknown rank and arbitrary position in the spectrum, while each retains its own distinct spiked directions. The model treats the two datasets symmetrically and strictly generalizes existing models for shared covariance structure. We develop a complete estimation procedure that includes joint estimation of the shared subspace and its rank, a closed-form pooling weight for combining the two datasets, and asymptotic guarantees derived from random matrix theory in the proportional-growth regime. The framework also resolves a gap in contrastive dimension reduction by providing a principled estimator for high-dimensional settings. We illustrate the methodology on portfolio construction during the early COVID-19 pandemic and on contrastive analysis of brain tumor gene expression.
Mufan Li, Jaume de Dios Pont, Mihai Nica +1math.PR stat.ML
We study the squared singular value spectrum of a non-square product of independent real Gaussian matrices, equivalently the feature covariance spectrum of a deep linear neural network at initialization. Starting from the fixed-$m$ covariance diffusion previously obtained in the proportional depth-width limit, we record an equivalent matrix realization, describe its affine invariance, and derive the interacting diffusion satisfied by its eigenvalues. We then take a second limit, sending $m\to\infty$ on the accelerated spectral clock $τ=mt$, which corresponds in this sequential construction to the relation $dm/n\to\barτ$. We establish convergence of the empirical spectral measure path to a deterministic mean-field limit and derive a closed Burgers equation for its $T$-transform. Together with the proportional depth-width limit, these results give a rigorous sequential route from the deep non-square Gaussian product to the free log-normal limit of its feature covariance spectrum; for more general initial laws, the transform yields a free multiplicative convolution form. We further analyze the support of the free log-normal law, give a fixed point iteration for numerical evaluation and a formal Marchenko--Pastur approximation at small time, and use the limiting spectrum to predict the risk in a toy random feature model.
Observable Matrix Dynamics (OMD) is a diagnostic framework that probes the dynamics of high-dimensional internal representations of inputs by a neural network via a fixed-size $N \times N$ distance matrix $M(t)$ on a held set of $N$ inputs. OMD uses methods of random matrix theory and particle dynamics to explore spectral reorganisations that are missed by scalar loss functions, but are informative of the training process. We read $M(t)$ against a perturbative ambient-versus-latent decomposition extending the Bogomolny--Bohigas--Schmit (BBS) theory of random distance matrices, with per-snapshot diagnostics for the top-of-spectrum band structure and ambient noise, trajectory-level observables linking snapshots, and a 3D MDS embedding (bottom-three eigenvectors) rendering training as a moving particle cloud. Across seven experiments, diffusive regimes lack stable top-of-spectrum band structure, while sharp endogenous or externally driven reorganisations produce stable fingerprints: consistent with smooth or product latent geometries in BBS-adjacent cases, and with finite-cluster or Fourier-soliton structures otherwise. OMD thus reads the geometric regime of a representation rather than reporting a single intrinsic dimension.
Fengkai Liu, Ke Wang, Wanjie Wangmath.ST cs.LG math.NA math.PR
Spectral methods rely on the stability of principal eigenspaces under random perturbations. Classically, this is quantified by the Davis-Kahan and Wedin theorems, which bound the eigenspace error via the operator norm of the noise and the relevant spectral gaps. While sharp for arbitrary deterministic perturbations, these worst-case bounds can be wasteful in the low-rank signal-plus-noise setting, as they fail to capture the interaction between the signal geometry and the noise distribution. We study the spectral perturbation of signal-plus-noise matrices corrupted by sparse random noise with an arbitrary, inhomogeneous variance profile. Under heterogeneous variances, the empirical eigenvectors suffer a systematic, deterministic geometric bias invisible to classical bounds. Leveraging the Quadratic Vector Equation (QVE) and fine-grained isotropic local laws, we derive near-optimal, non-asymptotic bounds for the leading eigenspaces in the operator and 2-to-infinity norms. These separate the usual signal-to-noise contribution, stochastic fluctuations, and structured geometric bias terms determined by the alignment between the signal eigenspaces and the row-wise variance profile. We further develop refined rowwise bounds that adapt to the variance-weighted leverage of the signal space, yielding sharper guarantees in delocalized regimes. As applications, we establish strong consistency of adjacency spectral clustering for degree-corrected stochastic block models with heterogeneous degrees and unbalanced communities, recovering the logarithmic expected-degree scale in the regular balanced case. We also study spectral embedding for generalized random dot product graphs, showing that the full signal embedding admits sharp rowwise control, whereas spectral truncation can retain a systematic geometric bias determined by the omitted signal directions and the variance profile.
Nicholas Barnfield, Juno Kim, Eshaan Nichani +2stat.ML cs.IT cs.LG
How many key-value associations can a $d\times d$ linear memory store? We show that the answer depends not only on the $d^2$ degrees of freedom in the memory matrix, but also on the retrieval criterion. In an isotropic Gaussian model for the stored pairs, we show that top-1 retrieval, where every signal must beat its largest distractor, requires the logarithmic model-size scale $d^2\asymp n\log n$. We prove that the correlation matrix memory construction, which stores associations by superposing key-target outer products, achieves this scale through a sharp phase transition, and that the same scaling is necessary for any linear memory. Thus the logarithm is the intrinsic extreme-value price of winner-take-all decoding. We next consider listwise retrieval, where the correct target need not be the unique top-scoring item but should remain among the strongest candidates. To formalize this regime, we propose the Tail-Average Margin (TAM), a convex upper-tail criterion that certifies inclusion of the correct target in a controlled candidate list. Under this listwise retrieval criterion, the capacity follows the quadratic scale $d^2\asymp n$. At load $n/d^2\toα$, we develop an exact asymptotic theory for the TAM empirical-risk minimizer through a two-parameter scalar variational principle. The theory has a rich phenomenology: in the ridgeless limit it yields a closed-form critical load separating satisfiable and unsatisfiable phases, and it predicts the limiting laws of true scores, competitor scores, margins, and percentile profiles. Finally, a small-tail extrapolation further leads to the conjectural sharp top-1 threshold $d^2\sim 2n\log n$.