The 95% Illusion: Why Your Confidence Interval Isn't What You Think It Is

| Source: Towards Data Science

Tags: statistics, confidence intervals, Bayesian inference, A/B testing, frequentist statistics, data science, experimentation

Most practitioners misinterpret 95% confidence intervals as '95% probability the parameter lies in this range' — but that is a Bayesian credible interval, not a frequentist CI. A CI describes the long-run hit rate of a procedure across repeated samples, not the probability attached to the one interval you computed, a distinction that routinely distorts A/B test decisions.

Details

Frequentist confidence intervals, introduced by Jerzy Neyman in the 1930s, are widely misread even by experienced researchers. The 95% figure is a property of the estimation procedure: if you ran the same experiment 100 times, roughly 95 of the resulting intervals would contain the true parameter. It says nothing about where the parameter falls given the specific interval you actually computed. The practical consequence surfaces constantly in product analytics. When a stakeholder asks whether there is a '95% chance the new version is better,' silence in the room signals a widespread misunderstanding. That probabilistic claim belongs to Bayesian credible intervals, which require specifying prior beliefs but produce posterior probability statements that actually answer the question. The article walks through Neyman's original motivation, the Central Limit Theorem mechanics behind CI construction, and a simulation of 100 intervals across repeated samples to illustrate what the 95% guarantee covers in practice. A credible interval does answer 'where is the parameter?' — but at the cost of requiring a prior. For ML and product teams running A/B tests, conflating the two frameworks leads to overconfident shipping decisions. The recommended fix is explicit: choose the inferential framework before analysis and match the statistical tool to the question being asked.