Skip to results
MLSift
← Feed
Statistical & Classical MLSemiparametric Inverse Problem2607.23130

Identification and Honest Recovery from Semantic Observation Kernels: Operator Error, Coarsening, and Stability

Matthew Francis Dixon

stat.ME stat.ML

Abstract

Probabilistic text generators, such as large language models, assign probabilities to phrases, but consequential decisions require posterior uncertainty over meaningful states. These are not interchangeable: language probabilities depend on the prompt, may be incomplete and need not reliably identify state uncertainty. Without a statistical bridge, fluent responses and numerical confidence are insufficient for inference or governance. We formulate recovery of the target posterior as a semiparametric inverse problem and develop honest recovery guarantees that account jointly for calibration error, measurement noise, incomplete probabilities and weak identification. Simulations demonstrate the predicted coverage and stability behaviour, while two frozen language-model studies demonstrate held-out recovery. The resulting method determines when a semantic measurement can be trusted for inference and when use, review, recalibration or abstention is warranted, providing a statistical foundation for runtime AI governance.

Topics

Classified with taxonomy v2 on Sat, 5 Sept 2026.

The PDF is 1–3 MB. Open it in your browser's viewer, or load it here.

Open PDF