Understanding the Limits of Agentic ICD Coding

| Source: arXiv AI

Tags: ICD-10-CM, medical coding, agentic AI, healthcare AI, MIMIC-IV, EMNLP 2026

A study accepted at EMNLP 2026 maps three distinct failure modes in AI-based ICD-10-CM medical coding: neural systems show a 0.43 micro-F1 gap on rare codes, workflow systems fail on injury/causality codes, and tool-augmented agents partially recover—but no single approach dominates across all conditions.

Details

ICD-10-CM coding is essential for medical billing and epidemiology, but standard benchmarks use aggregate metrics that obscure failures on specific code types. This EMNLP 2026 paper evaluates neural classifiers, workflow systems, and tool-augmented agentic systems on a rarity-stratified subset of MIMIC-IV discharge summaries.\n\nThe results expose two orthogonal failure modes. Neural classifiers show a 0.43 micro-F1 gap between rare and common codes—a quantified ceiling on reliability for infrequent diagnoses. Workflow systems handle rare codes but score near zero on injury and external cause codes that require multi-step guideline following.\n\nA tool-augmented agentic configuration with structured access to official ICD-10-CM reference materials recovers up to 0.34 micro-F1 on the injury/cause subset, but still no single system dominates across all conditions. The results suggest hybrid approaches may be necessary for production clinical coding rather than any single architecture.