Domain-Specific Jargon in Large Language Models: A Comparative Analysis between General-Purpose and Specialist Models

| Source: arXiv AI

Tags: Llama-3.1, fine-tuning, mechanistic interpretability, medical AI, domain adaptation, EMNLP 2026, jargon

Medical fine-tuning of Llama-3.1 counterintuitively hurts jargon comprehension versus the base model — mechanistic interpretability shows the fine-tuned model over-weights a small set of jargon-favoring components instead of redistributing parametric knowledge, accepted to EMNLP 2026.

Details

A widely held assumption in applied NLP is that domain fine-tuning improves specialized performance across the board, including on terminology. This paper challenges that assumption with a direct experiment: a Llama-3.1 model fine-tuned on medical data is evaluated against the base Llama-3.1 on two medical jargon benchmarks, and the base model wins on both. The authors explain this using mechanistic interpretability. Rather than reorganizing the model's parametric knowledge about medical terms, fine-tuning amplifies a small subset of model components associated with jargon-favoring predictions. This creates miscalibration: the fine-tuned model leans harder on specific attention heads and MLP layers in ways that hurt generalization. Component reweighting strategies — which suppress the identified jargon-overweighted components — close the performance gap between fine-tuned and base models on the benchmark tasks. Crucially, some of these jargon-sensitive components also activate for materials science jargon, suggesting they encode a partially domain-agnostic representation of specialized terminology rather than purely medical knowledge. Accepted to EMNLP 2026 main conference. The implications are practical: teams using domain fine-tuning to improve LLM performance on technical terminology should not assume they're getting better jargon comprehension — they may be making it worse. Mechanistic interpretability provides a diagnostic path to verify this.