Refusal Reads Only a Slice of What the Model Knows: Harm-Keyed Routing and Its Exceptions Across Model Families
| Source: arXiv AI
Tags: AI alignment, LLM safety, mechanistic interpretability, refusal, OLMo, Llama
Causal analysis across four open-weight model families shows LLM refusal reads only a single harm direction — not the broader moral reasoning subspace — with three-quarters of refusal's causal input lying outside moral content, explaining why single-direction edits bypass alignment while leaving moral comprehension intact.
Details
It has been known that a single direction can be edited out of a model's residual stream to disable refusal. What has been less clear is what the refusal gate was reading in the first place. This paper provides causal evidence across OLMo-3, Llama, Qwen, and GPT-OSS (four open-weight models, three families).\n\nThe central finding from OLMo-3: moral comprehension and the refusal gate are separate constructions. A low-rank moral subspace crystallizes during pretraining; alignment rotates it once without rebuilding it. The refusal gate, by contrast, is a fresh post-training construction built into a narrow control-token channel. A nested interchange rank sweep shows that as the basis widens, moral judgment reads progressively more of the moral subspace — but refusal levels off at the level of a single harm direction, with about three-quarters of its causal input lying outside moral content altogether.\n\nThe picture is not uniform across architectures: Llama reads broad moral content, Qwen extends beyond the single harm cue (unresolved at current sample size), and GPT-OSS reads harm with refusals that can be argued either direction by its own reasoning trace.\n\nThe implication for alignment is significant: refusal is routing on the harm percept, not on moral judgment. Where those diverge — or where adversarial inputs can isolate the harm channel — alignment can be bypassed without touching the model's moral knowledge.