Predictive coding offers a powerful theory of cortical computation, but corresponding scalable algorithmic implementations for artificial intelligence have remained elusive. This paper introduces the Bayesian reflex, a computational framework that directly instantiates predictive coding through three pillars: belief maintenance via hierarchical generative models, sequential Bayesian updating via prediction-error minimization, and uncertainty-driven action via active inference. We show that recent breakthroughs---ellipsoidal decomposition for exact $i.i.d.$ sampling, recursive Gaussian processes for deep hierarchical inference, and derivative-aware Bayesian optimization---provide the missing algorithmic ingredients. The resulting framework enables mathematically principled, scalable, and brain-inspired continual learning, perception, and decision-making. We illustrate its versatility through applications ranging from climate model evaluation to prime number discovery, offering a blueprint for truly adaptive artificial intelligence.
Hierarchical mixture models are a powerful tool for modeling data generated from heterogeneous sources, particularly when the mixing proportion $\boldsymbol{w}$ itself is treated as a random variable with a Dirichlet or Beta-Liouville prior. Such models are widely employed in scenarios where uncertainty in class membership or data-generating processes must be probabilistically quantified. This paper studies the exact marginalization of the mixture weight. For the two-component case we give an $O(n^2)$ dynamic program -- and an $O(n \log^2 n)$ FFT variant -- for the marginal likelihood, and show that the exact posterior of the weight is a finite mixture of Beta distributions, delivering closed-form posterior summaries, credible intervals and per-observation local false-discovery rates without any sampling. For $K \ge 3$ components we give an exact joint dynamic program. The gain is largest in the small-sample regime the method is built for: on a real multilevel meta-analysis, a pathway-level dysregulation analysis of leukemia gene expression, and a leukemia-derived gene-panel benchmark with known ground truth, the exact interval for the signal proportion is calibrated where EM gives no interval at all (collapsing to a boundary) and Gaussian/Laplace approximations mis-cover, and it is two orders of magnitude faster than the sampler that would match it. On the large prostate-cancer benchmark, where every method has ample data, it agrees with locfdr on the gene ranking while adding a posterior interval for the null proportion.
Qingyang Zhu, Eric Karl Oermann, Kyunghyun Chocs.LG
Bayesian predictive inference provides a principled framework for uncertainty quantification, data efficiency, and robust generalization. However, exact inference is often intractable, and scalable approximations may remain computationally expensive or require restrictive modeling assumptions that degrade predictive performance. Prior-Data Fitted and in-context models have recently emerged as an amortized alternative by learning to map datasets directly to predictive distributions, but existing approaches are tightly coupled to the support of the training prior and lack explicit mechanisms for adapting to new priors at test time, resulting in limited robustness under distribution shift. We introduce a multi-task in-context learning framework for amortized hierarchical Bayesian predictive inference that explicitly represents prior information as a prefix of in-context datasets. A transformer trained on sequences of prior and target tasks learns to adapt its predictions across families of priors. On a suite of evaluations with increasing difficulty, including out-of-meta-distribution priors and priors with high-dimensional latent structures, our method matches oracle Bayesian predictors while being orders of magnitude faster. We further demonstrate its practical relevance on a real-world spatiotemporal temperature prediction benchmark. Code is available at https://github.com/martianmartina/multi-task-bayesian-icl/.