LLMs Are Not (Consistently) Bayesian: Quantifying Internal (In)consistencies of LLMs’ Probabilistic Beliefs

| Source: Apple ML Research

Tags: Apple, Bayesian inference, LLM reasoning, probabilistic AI, AI reliability, Stanford, belief updating

Apple ML Research finds LLMs inconsistently update beliefs under new evidence — and counterintuitively, non-Bayesian heuristic updates often outperform exact Bayesian processing. The diagnosis: LLMs' underlying probabilistic world models are misspecified, not just their reasoning. Direct implications for AI deployed in medicine, science, and law.

Details

When LLMs are deployed in high-stakes domains — medicine, scientific reasoning, law — they must rationally update beliefs as new evidence arrives. Apple ML Research introduces the "information processing gap" as a formal measure of how far an LLM's belief update deviates from an optimal Bayesian update, then uses it to probe internal inconsistencies across multiple evidence-incorporation strategies. The core finding is counterintuitive: some tested approaches do produce near-Bayesian updates, but non-Bayesian heuristic updates frequently outperform exact Bayesian processing on downstream task accuracy. The researchers' interpretation is that LLMs' underlying probabilistic world models are misspecified — the internal representation of the world does not match reality closely enough for Bayes-optimal updating to help. The learned heuristic, by contrast, has absorbed implicit corrections for the prior's flaws, making it more practically useful even if theoretically suboptimal. The paper provides diagnostic methods to identify where LLM-powered inferential systems break down, offering a concrete toolkit for teams building reasoning chains, RAG pipelines, or clinical decision-support tools. Authors include researchers from Stanford and Apple's ML team. The practical upshot for practitioners: "make it more Bayesian" is not a reliable fix for unreliable LLM reasoning in high-stakes domains. World model misspecification is the deeper problem, and these diagnostics are a prerequisite before any remediation strategy can be designed.