Sumedh Gupte, Prashanth L. A., Sanjay P. Bhatstat.ML cs.LG
We consider the optimization of the Optimized Certainty Equivalent (OCE) risk, with applications including portfolio optimization in finance, and uncertainty quantification, classification, and regression in machine learning. Our contributions cover popular special cases of OCE, such as entropic risk, mean-variance risk, and smooth variants of Conditional Value-at-Risk. Our treatment sets out the conditions that facilitate the extension of OCE to unbounded r.v.s.. We provide a useful characterization of OCE that links OCE to utility-based shortfall risk (UBSR). Our characterization enables us to form an OCE estimator from the classic sample-average approximation (SAA) of UBSR. We derive mean-squared error (MSE) bounds for our proposed OCE estimator. For OCE optimization, we first derive an expression for the OCE gradient using the characterization linking OCE to UBSR. This expression serves as the basis for a gradient estimator for the OCE. We derive non-asymptotic bounds on the MSE for the proposed OCE gradient estimator. We incorporate the aforementioned gradient estimator into a stochastic gradient (SG) algorithm to optimize OCE and quantify its convergence rate using non-asymptotic bounds that we derive. Finally, we present three experiments that use our OCE optimization algorithm to solve portfolio optimization and uncertainty quantification problems.
We develop an approximate risk minimization framework for shrinkage-thresholding estimation in normal mean problems. In the canonical multivariate normal mean model, we introduce a general functional class of estimators that contains classical shrinkage and thresholding behavior, including James-Stein-type and lasso-type rules. We express quadratic risk as a functional over this class, derive optimality conditions for both oracle risk and data-driven approximate risk minimization, and construct a feasible approximate risk criterion from the observed data when the oracle risk is unavailable. The resulting estimator, NOMAD, is obtained by minimizing this approximate risk over the proposed class. For the canonical model, we develop an approximate risk minimization theory that includes optimizer characterization, sieve-based consistency under regularity conditions, and approximate-risk inequalities relative to benchmark procedures in the admissible class. We then extend the framework to multivariate normal mean estimation with correlated observations, develop both MLE-based and conditional MLE-based constructions, and establish consistency results under regularity conditions. We further apply the framework to linear regression and derive an equivalent penalized regression representation in which the shrinkage-thresholding map induces a data-adaptive penalty, recovering ridge-type and lasso-type behavior as special cases or limiting forms. The results provide a unified risk-based framework for shrinkage, thresholding, and regularization across canonical and correlated normal mean estimation and linear regression.
Mathieu Chalvidal, Florentin Coeurdoux, Eric Vanden-Eijndencs.LG stat.ML
We recast classical shrinkage of high-dimensional covariance estimators as empirical risk minimization over a parametric stochastic interpolant between a source and a target distribution. This formalism recovers known shrinkage estimators as special cases and reveals three distinct mechanisms for reducing statistical risk: (i) Scheduling: the interpolant schedule determines the class of admissible covariances, and hence the achievable risk. (ii) Flow maps and couplings: whereas naive constructions amount to assuming independence between the distributions, specific coupling structures (e.g., solutions of optimal transport problems) can lower the empirical risk. Moreover, non-linear flow maps realizing such couplings free the interpolant covariance from the eigenbasis of the empirical estimate, enabling eigenvector regularization. (iii) Early stopping: estimators defined by integrating a regressed vector field afford an additional bias-variance trade-off through approximation of the true interpolant distribution. We then propose a neural estimator of the interpolant, together with an upper bound on its quadratic risk in terms of the interpolant approximation error, and validate both on synthetic experiments. Finally, we apply the estimator to real neuroimaging data, demonstrating the additional regularization power this approach offers in practice.