NeuroActiSep: Detecting Factual Hallucinations from Feed-Forward Neurons in a Single Pass
| Source: arXiv AI
Tags: hallucination-detection, LLM-reliability, interpretability, mechanistic-interpretability, open-weight-models
NeuroActiSep identifies feed-forward neurons correlated with factual hallucinations at the final prompt token and shows their identities transfer across QA datasets — offering a single-pass hallucination detector that performs on par with probes trained on full internal state representations.
Details
Factual hallucinations in LLMs are expensive to detect — most reliable approaches require external retrieval, multiple model passes, or consistency checks against other outputs. NeuroActiSep takes a white-box approach: it ranks individual feed-forward neurons in an LLM at the final prompt token position based on their correlation with hallucination on a custom selection dataset, then trains lightweight probes using only those selected neurons' activations. The key finding is cross-dataset neuron transferability: probe identities selected on one dataset generalize to classify hallucinations on different factual QA benchmarks at performance comparable to probes trained on full internal state vectors. The paper also analyzes how the distribution of selected neurons varies across model layers, finding layer depth influences detection performance. The method requires white-box access to model internals — applicable to open-weight models deployed in-house, not API-only hosted models. For teams running their own Llama, Mistral, or similar models, this provides a computationally cheap single-pass hallucination signal without retrieval augmentation. Results are on factual QA benchmarks; generalization to instruction-following or multi-turn hallucination requires separate validation.