Behavioral contamination detectors can return "no evidence" either because a benchmark is clean or because the audit has little power. We formalize this distinction for a benchmark in which an unknown fraction alpha of items was seen during training. With matched clean and seen controls, the behavioral channel is the sparse mixture Q_alpha = (1 - alpha) P_0 + alpha P_1, and an exact second-moment argument shows that detectability is governed by alpha * rho * sqrt(m), where rho^2 = chi^2(P_1 || P_0) measures behavioral separability. Any scalar detector reduces to its efficacy, ef = |E_1 f - E_0 f| / sqrt(Var_0(f)) <= rho, which can be estimated from controls before the audit is run. A separate sample-split certificate lower-bounds alpha distribution-free, without requiring an orientation assumption. Our empirical finding is two-sided. Frozen calibration efficacy predicts held-out power curves, with R^2 = 0.83-0.98 across six exact-permutation channels, but the efficacy-only Gaussian budget is miscalibrated at the small sample sizes it prescribes, failing in 9/9 gate-passing channels even though efficacy itself transports. The failure is in the inversion, not the calibration. A predeclared two-stage planner that simulates the deployed test repairs the budgets, is uniformly conservative, and abstains when its probe does not transport. The certificate is valid but vacuous at audit scale, and a five-seed paired injection study recovers the mechanism ordering verbatim > paraphrase > surface, in which the apparent answer-only signal is explained by baseline drift. We report the audit contract and its failures together: a non-rejection is interpretable only alongside the efficacy, budget, and validity gates that produced it.
Divergence measures are essential tools for detecting distributional shifts in model monitoring, particularly crucial given the volatility of financial data. While the Population Stability Index is the most widely used measure, Jensen-Shannon Divergence and Kullback-Leibler Divergence offer distinct advantages. Jensen-Shannon Divergence handles mixture models, addresses zero-binning problems, and is symmetric, while Kullback-Leibler Divergence excels in Bayesian model comparison. This study extends the work of Yurdakul and Naranjo (2020) with two primary contributions. First, we derive the statistical properties and chi-square benchmark values for Jensen-Shannon Divergence and Kullback-Leibler Divergence. Second, we demonstrate their applicability by detecting distributional changes in credit default probabilities from Merton, Merton with jump, and stochastic volatility with jump models. Our results establish that Jensen-Shannon Divergence and Kullback-Leibler Divergence follow chi-square distributions and reveal important practical trade-offs. Jensen-Shannon Divergence exhibits superior Type I error control, maintaining rejection rates closest to 5%, thereby minimizing false positives. However, this conservatism reduces statistical power at small samples (27% versus 32% for Population Stability Index and Kullback-Leibler Divergence at n = m = 200), requiring larger samples for reliable detection. This trade-off enables practitioners to select measures based on whether minimizing false alarms or maximizing detection sensitivity is the priority.