Beyond Prompting: Efficient and Robust Contextual Biasing for Speech LLMs via Logit-Space Integration (LOGIC)

| Source: arXiv AI

Tags: speech recognition, Phi-4-MM, contextual biasing, speech LLM, multilingual, entity recognition

LOGIC injects entity probability boosts directly into the decoding layer of speech LLMs — achieving 9% relative reduction in Entity Word Error Rate across 11 multilingual locales with Phi-4-MM, while maintaining constant-time complexity regardless of entity list size.

Details

Customizing speech recognition for domain-specific terms — contact names, product names, technical jargon — is a perennial challenge for enterprise voice systems. Current Speech LLMs handle this via prompting, but as entity lists grow, prompting hits context window limits, increases latency, and suffers from 'lost in the middle' degradation. Post-processing via Generative Error Correction (GEC) introduces 'over-correction' hallucinations of entities never actually spoken. LOGIC (Logit-Space Integration for Contextual Biasing) operates directly in the decoding layer, injecting entity probability boosts into the logit space rather than the input prompt. This decouples context injection from input processing, achieving constant-time complexity regardless of entity list size — a critical scalability property for large enterprise deployments. Evaluated using Phi-4-MM across 11 multilingual locales, LOGIC achieves a 9% relative reduction in Entity WER with only a 0.30% increase in False Alarm Rate. This is v3 of the paper (originally submitted January 2026), suggesting the technique has been refined over several iterations. For enterprise deployments handling personalized assistants, call center transcription, or meeting notes where entity recognition accuracy matters, LOGIC represents a practical improvement over both prompting and GEC alternatives.