Many functional data analyses reduce random functions to scalar summaries or conditional mean curves. This is limiting when we wish to understand how covariates affect the distribution of entire functional responses, including their shape, timing, or variability. We study the problem of estimating conditional laws of functional outcomes and show that these objects can be estimated and evaluated in a practical nonparametric framework. To do this, we introduce functional distributional random forests, which estimate each conditional law as a covariate-dependent distribution over sampled functions by training a random forest to minimize a kernel-based maximum mean discrepancy within the leaf nodes of the decision tree. This supports inference on arbitrary functionals of the conditional distribution while keeping predictive samples tied to realistic curves. We consider a variety of kernels defined on function spaces, including Sobolev and operator-induced kernels. We also provide conditions for consistency of our estimator and develop scoring rules for comparing it to baseline estimators. In simulations, our method recovers distributional changes that are missed by baseline methods. In an application to NHANES accelerometer data, it identifies interesting covariate-associated changes in both median activity profiles and predictive dispersion.
We consider optimization problems defined on product spaces of simplices. Examples of this class of problems include learning low-rank discrete multivariate probability distributions via simplex constrained tensor decomposition and performing functional data registration under the Square Root Velocity Function (SRVF) representation. In this work, we demonstrate the feasibility of replacing the product simplex with a smooth, elementwise strictly convex reparameterization, resulting in an unconstrained optimization problem on a manifold. We show that performing such a reparameterization results in the second order Karush-Kuhn-Tucker (KKT) points on the smooth manifold being mapped to the weak second order KKT points on the product simplex. This leads to a Riemannian Gradient Descent (RGD) algorithm for solving the reparameterized problem, which outperforms Projected Gradient Descent (PGD), and provides a more faithful representation of the original function shapes while performing curve registration.
Anomalies in functional data arise from rare or distinct processes that deviate from the dominant data-generating mechanism. Detecting such departures is essential in applications where they may correspond to errors, structural changes, or other behavior of interest. This work introduces a Bayesian nonparametric approach for anomaly detection in multivariate functional data. We model functional data as an infinite mixture of multi-output Gaussian processes, with a finite and automatically determined number of mixture components obtained through slice sampling. Mean functions are represented using a wavelet basis and regularized through Besov priors to obtain a smooth and sparse representation of the data. Cross-functional dependence is captured using the intrinsic coregionalization model and we solve covariance kernel selection by introducing a Carlin-Chib product space step in the Markov Chain Monte Carlo algorithm. Within this model, anomalous observations are assigned to small mixture components without requiring prior specification of the number or nature of anomalies. We consider a semi-supervised setting, in which labels are available for 15% of the normal observations and a large class imbalance is present. The utility of our model is demonstrated on both univariate and multivariate functional data.
Aleix Alcacer, Rafael Benitez, Vicente J. Bolos +1stat.ME cs.LG stat.AP
We introduce biarchetype analysis for the first time in the context of univariate functional data. This unsupervised methodology extends archetype analysis by simultaneously identifying archetypal structures across both the cases (countries, in our application) and the temporal argument. Both cases and time points are expressed as mixtures of biarchetypes, yielding a concise and highly interpretable representation of complex functional observations. Although biarchetype analysis is not intended as a clustering technique, it offers superior interpretability compared with biclustering approaches, as it is based on extreme, representative patterns rather than average centroids, thereby enhancing human comprehension. We apply the proposed method to 10-year government bond yields of European countries over the period 2001-2025. The results identify three distinct time regimes (the pre-crisis period, the euro-area sovereign debt crisis, and the post-crisis period), and reveal Germany, Greece, and Hungary as country archetypes.