Let $X(t)$, $t\in K$, be a centred Gaussian process with continuous sample paths on a compact metric space $K$, and let $M=\min_{t\in K}X(t)$. Let $σ_*^2$ denote the minimum covariance energy associated with $X$, and assume that $σ_*^2>0$. Motivated by the results of \cite{chakrabarty2018asymptotic} for smooth Gaussian processes, we show that, conditionally on $M>u$, the scaled overshoot $u(M-u)$ converges, as $u\to\infty$, to an exponential random variable with mean $σ_*^2$. Moreover, every weak subsequential limit of the conditional law of a measurable minimizer of $X$ is an optimal covariance-energy measure. In particular, if this measure is unique, then the conditional law converges weakly to it. The results are illustrated by stationary Gaussian processes, fractional Brownian motion, and fractional Brownian sheet.
We introduce Self-Similar Generative Estimation (SS-GEN), a method for simulating multivariate tail events and estimating rare-event probabilities in both heavy and light-tailed settings. SS-GEN exploits asymptotic tail structure to decompose the tail distribution into an explicit radial component and a nonparametric angular component, reducing tail learning to a compact-domain problem that can be handled by off-the-shelf deep generative models. The resulting sampler generates representative extreme scenarios and supports probability estimation far beyond the observed data. Under mild nonparametric tail assumptions, we show that the SS-GEN density is asymptotically exact in the tail, with vanishing uniform relative error for regularly varying distributions and vanishing uniform log-relative error for Weibull-type distributions. Unlike existing approaches that rely on specialized architectures or parametric tail specifications, SS-GEN leverages asymptotic tail structure to enable standard generative models to generate representative extreme samples and estimate rare-event probabilities beyond the observed data.
We develop a statistical learning theory for gradient boosting applied to the estimation of covariate-dependent Generalized Pareto (GP) distributions in the context of Peaks-over-Threshold modeling. After an orthogonal reparametrization of the GP likelihood that diagonalizes its Fisher information matrix, we cast the estimation problem within the Empirical Risk Minimization (ERM) framework and derive non-asymptotic error bounds for the boosting estimator. Our analysis accounts for three distinct sources of error in the process: statistical fluctuations, the approximation bias inherent to the asymptotic nature of the GP model-controlled under second-order regular variation-and the approximation error associated with the finite number of boosting iterates, making explicit the resulting bias-variance trade-off. We illustrate the practical benefits of the reparametrization through simulations, showing that it significantly reduces gradient correlation during training and improves convergence stability. The methodology is applied to a medical malpractice insurance dataset from the Texas Department of Insurance, comprising over 18 000 closed claims. The gradient boosting approach yields a good fit for the tail of settlement cost distributions and reveals that the number of days to settlement is the dominant predictor of tail heaviness, consistent with earlier findings in the reserving literature.
Dan Cooley, Anne Sabourin, Troy Wixsonstat.ME math.ST stat.ML
This chapter explores ways to reduce the dimensionality of the data while preserving key information relevant to the analysis of multivariate extreme values.