Lost in Phonation: Voice Quality Variation as an Evaluation Dimension for Speech Foundation Models
| Source: arXiv AI
Tags: speech AI, bias, paralinguistic, foundation models, fairness, Interspeech 2026
VQ-Bench reveals that leading speech foundation models systematically shift responses based on voice quality — attributing different agency, empathy, and leadership to modal vs. breathy vs. creaky voices — with gender asymmetries in salary and leadership endorsements; accepted at Interspeech 2026.
Details
Standard speech AI evaluations focus on transcription and emotion recognition, leaving paralinguistic variation — voice quality like breathiness or creakiness — entirely untested. VQ-Bench fills this gap with a controlled parallel dataset of synthesized modal, breathy, creaky, and end-creak phonation types. Authors evaluated several speech foundation models on open-ended generation across four ecologically valid domains plus speech emotion recognition. Results show systematic shifts: models attribute higher agency, leadership, and empathy to some phonation types over others. A leading commercial speech API failed basic biometric sanity checks — it could not maintain consistent identity assumptions across voice quality variations. Most concerning: gender asymmetries emerged in salary and leadership endorsements. Models appeared to mirror or amplify human social biases when responding to gendered voice qualities — a bias vector invisible in standard benchmarks. Accepted at Interspeech 2026, the field's leading conference for speech AI.