IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows

| Source: arXiv AI

Tags: voice agents, benchmarks, conversational AI, enterprise AI, speech models, interruption handling

IHBench reveals that closed-weight voice models (OpenAI, Google) are 3.3x more robust to mid-conversation interruptions than open-weight alternatives, and that post-interruption recovery quality is a distinct evaluation axis not captured by any existing speech benchmark.

Details

Voice agents in customer service, healthcare scheduling, and account management face frequent user interruptions. Current benchmarks measure when interruptions happen (barge-in timing, turn-taking) but not what happens afterward — does the agent resume at the correct workflow step, address the user's interjection, and avoid repeating already-delivered content? IHBench evaluates 27 audio-language model configurations from OpenAI, Google, and the open-weight community across 10 enterprise domains. Six interruption types are injected at controlled mid-utterance points, scored on two axes: task fulfillment and recovery quality. The headline findings: closed-weight models (from OpenAI and Google) consistently outperform open-weight alternatives, degrade 3.3x more slowly as conversation length grows, and show no audio-versus-text modality gap. Open-weight models show a modality gap and underperform on all three dimensions. Cross-benchmark analysis against AudioMultiChallenge confirms recovery quality is a distinct capability axis not captured by existing evals. A human study validates the LLM judge against human annotators. For enterprise teams evaluating voice AI deployments, IHBench provides a concrete differentiating capability dimension. Authors include Alex Smola (AWS).