Language models are compared by their held-out per-token cross-entropy risk---the quantity scaling laws are fitted to. We show that it cannot be consistently estimated. Consistency, or convergence to the estimand, is defined relative to a \emph{possible state of the world}: a pair consisting of a data-generating distribution and a model we turn out to train. Quantifying over models as well as data-generating mechanisms is essential, because what decides whether a model's risk is estimable is a tail property of the distribution its weights induce, which no sample reveals. The per-token cross-entropy risk is hard to estimate because of a topological fact: among the possible states, finite risk and infinite risk each lie arbitrarily close to every instance of the other. Consequently no estimator---not merely the holdout average---is consistent at every state at which the risk is defined. Worse, inconsistent estimation persists under both bounding the expected sequence length and restricting to full-support models; and in that restricted setting the states at which inconsistency occurs are even dense. Two interesting ways out are identified, and neither is free. Way out 1: using a bounded context window, we can floor a model's next-token probabilities, making its risk finite exactly when the data-generating distribution has finite expected sequence length---a new, statistical rationale for a choice that was made on computational grounds, though the assumption it substitutes is itself beyond the reach of any test. Way out 2: reporting the risk only when it falls below a threshold fixed in advance restores consistency, at no cost to what model selection actually requires---but we need to recognize that the goal of estimation is revised.
As generative artificial intelligence enters scientific and professional work, its uncertainty must be defined on the states that matter for inference and decision-making. Language models assign probabilities to words, whereas applications require uncertainty over meaningful states such as diagnoses, hypotheses or operational conditions. We introduce a \emph{semantic map}: a prespecified, testable bridge from probabilities over verbal responses to a posterior over declared finite states. The language distribution remains unrestricted; held-out calibration connects it to a reference posterior. We derive posterior-error bounds and conditions for existence, conditional uniqueness, presentation stability and stable inverse recovery. This distinction matters because language probabilities depend on prompt wording, while the target posterior should not change under information-equivalent rewording. Experiments use professional market text compiled from Federal Reserve economic and financial series, together with controlled simulations having exact posteriors. Across two fitted language models, language-derived probabilities outperform printed numerical confidence, recover held-out posteriors with valid uncertainty coverage, remain largely stable under paraphrase and respond appropriately to altered evidence. \textbf{Prompt engineering optimises a wording-dependent response; robust scientific use requires validated stability of application-relevant meaning.} The proposed map turns semantic uncertainty in generative systems into an identifiable and testable statistical measurement problem and, when its acceptance conditions hold, yields an auditable posterior estimate.
Large language models exhibit remarkable emergent behaviors, yet the physical mechanism governing their collective dynamics remains poorly understood. Cognitive Field Theory predicts that learning reorganizes the collective relaxation spectrum, thereby modifying memory self-energy, long-memory dynamics, and collective susceptibility through the infrared organization of slow relaxation modes. Here we test this framework directly in Transformer dynamics. Using publicly available Pythia language models, we extract relaxation spectra from layer Jacobians throughout training, prompt ensembles, network depth, and model scale, allowing the collective observables of Cognitive Field Theory to be measured quantitatively. The measurements reveal pronounced infrared reorganization of the relaxation spectrum. Learning substantially redistributes spectral weight while preserving a nearly flat but weakly infrared-enhanced time-scale density of states, \( ρ(λ)\simλ^β, \qquad β\simeq-0.1, \) with a corresponding memory kernel exhibiting robust long-memory scaling close to \( K(t)\sim\frac{1}{t}. \) The collective observables further reveal a critical formation process: the memory self-energy reaches a transient maximum during early training before relaxing toward a metastable near-critical regime. Prompt-resolved and token-subspace measurements show that distinct local Jacobians recover a common macroscopic TDOS with shared infrared scaling, consistent with infrared fixed-point organization under coarse graining. The reproducibility of this infrared organization across training, prompt ensembles, network depth, and Transformer model scales supports infrared slow-mode organization as a robust collective principle of Transformer dynamics, providing a quantitative experimental realization of the collective observables predicted by Cognitive Field Theory.