Kai Yang, Masoud Asgharian, Celia M. T. Greenwoodmath.OC cs.AI math.ST stat.CO stat.ML
This paper addresses the limitations of Gaussian distribution assumptions in statistical sparse learning, particularly in modeling correlated and heterogeneous data. Conventional Gaussian models often lack robustness towards outliers and underlying distribution assumptions. To overcome these limitations, we propose the use of the $q$Gaussian distribution, derived from Tsallis entropy maximization, as a robust alternative. This is notably relevant in biostatistics, where the presence of correlated observations and heterogeneity, such as in genetic and longitudinal studies, is prevalent. Our contributions include modeling of correlated data through the re-derived multivariate probability density function from Tsallis entropy maximization, thereby addressing the limitations inherent in conventional Gaussian models. Furthermore, we introduce a novel framework that adapts numerical methods designed to find equilibria in flows to tackle composite optimization problems prevalent in statistical sparse learning. Applying this framework to the Hager-Zhang conjugate gradient algorithm \cite{Hager2005}, we develop a numerically stable and efficient algorithm for sparse statistical learning. The $q$Gaussian distribution, informed by the principle of maximizing Tsallis entropy, presents a viable and flexible alternative to Gaussian-based methods. This paper not only contributes to the theoretical understanding of statistical distributions and optimization techniques, but also paves the way for practical data analysis.
Graph attention normalizes neighborhood scores with softmax, the maximum-entropy choice under Shannon statistics. But homophilic and heterophilic graphs want different attention shapes, and one fixed normalization cannot serve both. We propose \textbf{LTGA} (\textbf{L}earnable \textbf{T}sallis \textbf{G}raph \textbf{A}ttention), a graph attention layer whose Tsallis entropic index $q$ is learned jointly with the weights, interpolating continuously between heavy-tailed ($q\!<\!1$), softmax ($q\!=\!1$) and compact-support ($q\!>\!1$) attention at four granularities from a global scalar to a per-edge index, under a bounded reparameterization that starts every model at the GAT baseline. Across eight benchmarks at ten seeds, LTGA-Edge takes the best average rank ($2.75$), but the omnibus test does not reject ($p\!=\!0.199$) and learning $q$ does not beat searching it: a validation-tuned frozen grid reaches $61.4\%$, tuned $α$-entmax $62.2\%$ and a capacity-matched $q\!\equiv\!1$ control $62.0\%$, against $61.7\%$ for LTGA-Edge. What the learned index buys is one run instead of a grid, and an interpretable mechanism: where $q$ leaves $1$, it prunes $42\%$ of attention coefficients to exactly zero, and those edges are selectively the wrong ones, restoring them costs $7.1$ points, while random pruning at the same rate costs $13.0$ more. Project page: https://kleyt0n.github.io/ltga
Maximum-entropy reference distributions are usually constructed on the normalized probability simplex. This formulation is less natural for unnormalized statistical models, in which positive multiples represent the same shape, and it does not directly explain how a prescribed admissible region should determine the deformation parameter of a bounded-support reference distribution. We formulate maximum entropy on the projective space of nonnegative measures and establish three results of statistical relevance. First, a universality theorem shows that every admissible monotone transform of the same normalized power functional has exactly the same optimizer under linear moment constraints. The result unifies the maximum-entropy implications of Tsallis and Rényi entropies, Hölder composite scores, pseudo-spherical scores, Bregman--Hölder constructions, and related homogeneous divergences without asserting a new distribution family. Second, the common optimizer is characterized as a $q$-exponential density; under mean and covariance constraints it is a compactly supported $q$-Gaussian for positive deformation and a Student-type density for negative deformation. Third, a prescribed Mahalanobis acceptance region with squared radius $R^2>d+2$ uniquely determines the deformation parameter $γ_R=2/(R^2-d-2)$. The resulting affine-equivariant reference density is the unique projective maximum-entropy solution, and its support coincides with the specified ellipsoid without an additional support constraint. This provides a principled method for constructing bounded-support statistical reference distributions from robust location and scatter estimates or from externally specified admissible regions.