AI models' written reasoning steps correspond to distinct internal patterns, a new study finds
| Source: THE DECODER
Tags: interpretability, chain of thought, Qwen, Gemma, KAIST, reasoning, mechanistic interpretability
Researchers at KAIST and Naver AI Lab found that eight distinct reasoning operations — extraction, decomposition, formula retrieval, deduction, computation — are reliably separable in LLMs' internal activations, with the clearest signal in middle layers, holding across Qwen2.5-7B, Qwen3-8B, and Gemma4-31B.
Details
The study asks whether the reasoning steps a model writes out (read data, break down problem, retrieve formula, compute) are also distinguishable in the model's internal representations — or whether the chain of thought is purely surface-level text with no internal analog.\n\nThe team from KAIST and Naver AI Lab defined eight reasoning operations and had three models (Qwen2.5-7B, Qwen3-8B, Gemma4-31B) solve math problems. GPT-5 labeled each segment of the solution path with one of the eight operation types. They then probed whether a classifier could tell reasoning steps apart from activation patterns alone.\n\nKey findings: (1) Reasoning operations are reliably separable across all three models, with the clearest signal in middle layers. (2) A token-vocabulary-only classifier underperformed an activation-based one — meaning internal states carry information beyond word choice. (3) Common function words ('a', 'is', 'the') cluster separately by early vs later layers depending on which reasoning step they belong to — the same word gets different representations in different reasoning contexts. (4) Blocking attention to the preceding 30 tokens via targeted intervention weakened the reasoning step signal, showing steps build on context rather than emerging in isolation. (5) The reasoning-type signal persisted even on incorrectly solved problems.\n\nThis matters for AI safety and interpretability: models process more internal structure than their visible chain of thought reveals, and that structure can potentially be read from activations.