Not All Speech Is Intent: Adaptive Self-Correcting Inference Layer for Post-ASR False Wake-Up

| Source: arXiv AI

Tags: wake word detection, ASR, voice AI, conversational AI, false wake-up, online learning, personalization

ASCIL is a post-ASR correction framework that cuts false wake-up errors by 54.27% relative on a session-disjoint subset, using acoustic embeddings, hesitation/silence signals, and personalized learning from past misclassifications — with under 60ms added latency.

Details

False wake-up activations — where a voice assistant triggers on speech that was not directed at it — remain a persistent UX problem in conversational AI. Phonetically similar speech can produce valid ASR transcripts that the assistant incorrectly executes.\n\nASCIL (Feedback-Driven Adaptive Self-Correcting Inference Layer) is a complementary post-ASR layer that re-evaluates wake-up intent before response generation. It fuses four signal types: acoustic embeddings, linguistic cues, device context (is the user looking at the screen? is there background noise?), and patterns from historical misclassifications for that user.\n\nCritically, ASCIL updates its patterns online from implicit behavioral signals (hesitation, silence, disengagement, cancellation) without requiring manual annotation. The intentional/unintentional labels used for evaluation are human-annotated, but ongoing pattern learning uses noisy inferred signals.\n\nEvaluated on 3,667 human-annotated interactions spanning 14 acoustic and contextual conditions, ASCIL achieves 54.27% relative error reduction on a session-disjoint failure subset. At threshold 0.90, it achieves 24.39% relative error reduction on the issue-tagged evaluation slice, while improving intentional acceptance rates. Median added latency: under 60ms.\n\nThe dataset is proprietary. This is a research paper from the ASR/conversational AI space, not a product announcement.