Hindsight Bias in Clinical Temporal Reasoning: How Future Data Exposure Affects Large Language Model Judgment

| Source: arXiv AI

Tags: clinical AI, hindsight bias, healthcare, temporal reasoning, GPT, benchmark, ML4H

A new benchmark of 171 clinical case reports shows LLMs including GPT, Gemma, and Opus consistently fall into hindsight traps when given full patient timelines — rating past decisions with knowledge of outcomes they couldn't have had at the time — and temporal record masking measurably reduces this bias.

Details

Clinical decision-making is prospective, but most clinical LLM evaluations use retrospective records that reveal the final diagnosis and outcome. This paper introduces a paired benchmark to measure how much LLMs' judgements shift when they can see the full timeline versus only information available at a clinical cutoff. The benchmark covers 171 PubMed Central case reports: 40 sepsis cases and 131 GLP-1/diabetes cases, represented as both narrative text and textual time series (TTS) with human-annotated and LLM-generated annotations. For each case, questions are anchored to a clinical cutoff and paired with a prospective reference answer and a hindsight trap answer consistent with the final outcome. Models tested: GPT 5.6 Sol, Gemma 4, GLM 5.2, and Opus 5. All four show consistent hindsight-sensitive shifts with full timeline exposure. Temporal masking — showing only information available at the cutoff — reduces bias without lowering accuracy on prospective questions. The paper was accepted at ML4H 2026. The implications are practical: any clinical LLM evaluated on retrospective records may be rewarded for using future information it would not have at decision time. This is a methodological warning for clinical AI validation as much as a model finding.