Yeari Vigder, Paulina Hoyos, David Thong +3cs.LG math.NA math.ST
Spectral embedding methods are widely used for dimensionality reduction and clustering of high-dimensional datasets with intrinsic low-dimensional structures. Although many datasets of practical interest exhibit invariance under symmetries such as rotations, standard spectral embedding methods do not account for this, treating symmetry-related data points as unrelated. Our approach to this problem is to incorporate the symmetries directly into the affinity kernels used for spectral embedding. We analyze the case of a Riemannian data manifold $M$ with symmetries given by a compact Lie group~$G$ and prove that, under suitable conditions, graph Laplacians constructed from three types of invariant kernels converge pointwise to explicit second-order differential operators on the quotient space $M/G$. Our analysis implies improved convergence rates, as the effective dimension drops according to the dimension of the group. We validate our approach on datasets with $\mathrm{SO}(2)$ or $\mathrm{SO}(3)$ symmetry, and show that $G$-invariant spectral embedding recovers the intrinsic geometry of the data, in contrast to standard spectral embedding, which fails to do so even in the limit of infinite data.
When analyzing a manifold learning algorithm for data lying on a smooth, compact, connected Riemannian submanifold $(\mathcal{M}, g)$ of $\mathbb{R}^d$, a key estimate for the geodesic distance $d_g$ is that there exists $K > 0$ such that $0 \leq d_g(p, q)^2 - \|p-q\|^2 \leq K d_g(p, q)^4$ for all $p, q \in \mathcal{M}$. We observe that more generally, when $\mathcal{M}$ is equipped with a smooth symmetric divergence $D$ satisfying a non-degeneracy condition and $g$ is given by $g_p := \frac{1}{2}\mathrm{Hess}_p(D(p, \cdot))$ for all $p \in \mathcal{M}$, there exists $K > 0$ such that $\left| D(p, q) - d_g(p, q)^2 \right| \leq K d_g(p, q)^4$ for all $p, q \in \mathcal{M}$. We demonstrate that this is sufficient for the pointwise convergence of graph Laplacians constructed with $D$ and discuss examples where $D$ is given by the Sinkhorn divergence on a family of probability measures parametrized by a manifold.
Ignacio Echave-Sustaeta Rodríguez, Aida Abiad, Frank Röttgerstat.ME stat.ML
Graph Laplacians encode graph structures in matrix form, and thus facilitate the application of linear algebra to graph theory. In statistics, two related families of probabilistic graphical models can be parameterized by graph Laplacians. The first one is the Laplacian-constrained Gaussian graphical model (LCGGM), which imposes that the (pseudo-)inverse covariance matrix of a Gaussian random vector is a Laplacian matrix. Applications include graph signal processing and network topology learning. The second one is the Hüsler-Reiss graphical model, which is considered as an extremal analog of the Gaussian graphical model, and can be used in extremal dependence modeling of floods, heatwaves, and financial losses. For both models, the restriction to positive edge weights in the graph Laplacian gives rise to an approach for graph structure learning that does not require tuning parameters. While these approaches yield a strong model fit in many settings, the resulting graph estimates are typically much denser than the underlying ground truth, limiting interpretability and scalability. In order to improve the accuracy of Laplacian-constrained graph learning, we propose to use spectral graph sparsification as a post-estimation operation. To do so, we replace the original Laplacian estimate by a sparser Laplacian that is spectrally close, and re-fit the model on the resulting graph. We refer to the two resulting methods as Spectral-LCGGM and Spectral-HR. We investigate the properties of the proposed estimators and show several theoretical results on their performance. Furthermore, we demonstrate that the newly proposed methods perform well by running simulations on Erdős-Rényi and stochastic block model graphs, and we also showcase their applications to real data.