Babak F. Dehkordi, Jeffrey L. Andrews, Andrew Jirasekstat.ME stat.ML
We propose a multivariate extension of the pseudo-Voigt profile-a weighted convex combination of Gaussian and Cauchy distributions-within a finite mixture modeling framework for robust model-based clustering and outlier detection. To ensure parsimony and coherence within clusters, shared location and scale parameters are imposed between the Gaussian and Cauchy components. Parameter estimation is carried out via an Expectation Maximization algorithm, with latent variables facilitating efficient likelihood-based inference. The performance of the proposed model is evaluated through simulation studies and applications to real-world data. Comparisons with established robust models, including mixtures of contaminated normal distributions, are provided to illustrate the model's clustering accuracy and outlier detection capabilities. The framework is shown to be particularly effective for data characterized by heavy-tailed behavior.
Given a dataset where a portion of the samples are contaminated, our goal is to recover the underlying clean population distribution. To this end, we propose Wasserstein Filtering (WF), a novel sample selection framework that discards a fraction of suspicious samples and estimates the target distribution using the empirical measure of the remaining data. The core insight is to select a subset of samples whose empirical distribution maximizes its Wasserstein distance to the fully contaminated empirical distribution, thereby preferentially isolating and removing geometrically influential outliers. To render this optimization computationally tractable, we introduce three algorithms: a marginal screening scheme, SinkMarg, and two joint optimization algorithms, SinkWF and SlicedWF, leveraging entropic optimal transport and sliced Wasserstein approximations, respectively. On the theoretical front, we introduce the Far Exclusion and Local Projection (FELP) contamination model, which characterizes corruptions consisting of well-separated outliers and locally indistinguishable perturbations. Under this model, we prove that the WF estimator achieves minimax optimality over distribution families with bounded covariance. Extensive numerical experiments on synthetic datasets, benchmark anomaly detection suites, and robust generative learning with diffusion models demonstrate that WF serves as a highly practical, model-agnostic preprocessing tool. It delivers competitive outlier detection performance and provides substantial downstream benefits for generative modeling under heavy contamination.
Generalised Bayesian inference (GBI) has emerged as a compelling robust alternative to standard Bayesian inference, mitigating sensitivity to data contamination by replacing the log-likelihood with a robust loss or divergence. However, existing robust GBI frameworks typically provide only qualitative robustness: while they can make posterior inference less sensitive to contamination, they lack an intrinsic mechanism to quantify the contamination proportion or identify anomalous observations. This paper introduces Hölder-Bayes, a GBI framework for joint inference of the model parameter and the contamination proportion. We construct a generalised joint posterior over both model and contamination parameter by applying the Hölder divergence to a scaled model density. Theoretically, we establish global bias-robustness via the uniform boundedness of the posterior influence function, derive a finite-sample excess-risk bound, and prove a Bernstein--von Mises approximation together with interpretable contamination-induced bias bounds under a heavy-contamination regime. We further show that, for the Hölder posterior, temperature calibration admits a direct interpretation as affine volume scaling of the data space. The resulting posterior yields a self-contained probabilistic mechanism for outlier detection: posterior uncertainty in both the model parameter and the contamination proportion is propagated to observation-level Frequency-of-Detection scores, without requiring an external anomaly-score threshold. Empirical evaluations demonstrate that Hölder-Bayes provides robust parameter inference, contamination-level recovery, and uncertainty-aware outlier detection.
William Roy Orchard, Philipp M. Faller, Dominik Janzingstat.ML cs.LG
True causal relationships are rarely known, and inferring causal graphs from data is hard. A fundamental challenge is how to assess whether a given causal graph is good in the absence of a ground truth. We propose falsifying candidate causal graphs based on whether they can explain the propagation of an outlier event. Our approach leverages a key principle: weak outliers rarely cause strong ones. While this principle has previously been used in root cause analysis to identify root causes without prior knowledge of the graph, we turn it on its head and use it to falsify candidate causal graphs whose implied outlier propagation is inconsistent with the data. To this end, we present the first statistical tests for the hypothesis that a candidate graph is the true causal graph, and show they have false positive control, power guarantees against incorrect causal graphs, and can operate with a single outlier sample.
Outlier detection (OD) aims to identify anomalous instances by learning the underlying structure of normal data (inliers), and is particularly challenging in fully unsupervised settings where no information about anomalies is available during training. Recent advances have leveraged the inlier-memorization (IM) effect, a phenomenon in which deep models memorize inlier patterns earlier than those of outliers, as a powerful signal for distinguishing outliers. However, despite its empirical success, the theoretical understanding of the IM effect remains limited. In this work, we present a theoretical study of the IM effect. Focusing on a simple autoencoder, we show that, under mild assumptions, the model can successfully memorize inliers while failing to memorize outliers during certain stages of early training. In particular, we characterize not only the emergence of the IM effect, but also its strength and persistence, and analyze how these properties depend on the data distribution and parameter initialization. In addition, building on these insights, we derive simple yet practical guidelines for enhancing the IM effect, including data preprocessing and parameter initialization schemes, achieving state-of-the-art performance on the ADBench datasets. Our findings provide a theoretical foundation for the IM effect and offer actionable directions for improving IM-based outlier detection methods.
Catarina P. Loureiro, M. Rosário Oliveira, Paula Brito +1stat.ME stat.ML
Explainability is increasingly recognized as a key aspect of outlier detection. However, for complex data structures such as interval-valued data, it remains largely unexplored. Building on an outlier detection framework based on the Interval Minimum Covariance Determinant estimator, we propose a novel approach to explain the outlyingness of interval-valued observations using the concept of the Shapley value. We derive a closed-form expression for the Shapley value of the squared robust Interval-Mahalanobis distance, enabling efficient computation of variable contributions. This formulation allows for a fine-grained interpretation of outliers, providing a detailed decomposition into contributions from centers, ranges, and cross-terms of the interval-valued observations. Moreover, the Shapley value is closely connected to the concept of cellwise outliers, as it can help identify variable-specific outliers that may not be evident at multivariate level. We further extend the framework through the Shapley interaction index to capture pairwise variable interactions driving atypical behavior. The practical utility of the proposed approach is illustrated through two real-world datasets.
Standard NLP pipelines for occupational clustering discard the 10-15% of job postings that density-based methods assign to noise. We argue this is an error: in rapidly evolving domains, low posting density signals novelty, not incoherence. We formalize this as the Emergence-Density Inversion (EDI) hypothesis and test it longitudinally on 84,988 job postings across eight quarters (Q4 2022-Q3 2024). EDI is partially confirmed: high-EOS outlier groups transition to stable clusters in 1.4 +/- 0.6 quarters vs. 4.1 +/- 1.2 for low-EOS groups (p < 0.001), though the signal fails in approximately 19% of cases, which we characterize as a failure analysis. We extend the Emerging Occupation Score (EOS) with Temporal Velocity and Cross-Platform Convergence, improving 2-quarter cluster-formation prediction from F1 = 0.61 to 0.74, outperforming Isolation Forest, LOF, GLOSH, and BERTrend baselines. A retrospective study on three now-established roles (MLOps Engineer, DevOps/SRE, Data Engineer) confirms EOS signalled 2-3 quarters before cluster formation, providing held-out validation. A held-out annotator panel (kappa = 0.74) rates EOS > 0.75 as coherent emerging occupations with 77% precision. Prompt Engineer, AI Safety Researcher, Foundation Model Engineer, and Agent Systems Engineer, all absent from O*NET, are top-4 in Q3 2024 and form stable clusters by Q1 2025.
Arthur Hendricks Mendes de Oliveira, Giovani Valdrighi, Marcos Medeiros Raimundocs.LG
The increasing use of machine learning algorithms in social applications has raised concerns about fairness and transparency, leading to the development of counterfactual explanations. These explanations supports individuals to understand and potentially alter unfavorable decisions in areas such as loan applications, job selections, and more, by providing actionable changes to input features that would lead to a desired outcome. Existing methods often struggle to balance feasibility, plausibility, and computational efficiency. To address this, we introduce P$^2$CE, an algorithm for generating plausible Pareto-optimal counterfactual explanations, offering users a diverse set of optimal trade-offs between different notions of feasibility. P$^2$CE employs an auxiliary isolation forest outlier detector to ensure that explanations are in accordance with the data distribution and leverages SHAP values to obtain optimal results with short computing times, regardless of the underlying model. Our algorithm was empirically evaluated on three datasets, demonstrating superior performance in terms of both solution quality and computational efficiency compared to related techniques.