Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
| Source: arXiv AI
Tags: speech-anonymization, privacy, F3-VA, voice-privacy, speech-AI, GDPR, data-compliance
A two-stage speech anonymization framework replaces identifiable content via generative editing and applies flow-matching-based voice anonymization (F3-VA) — maintaining ASR, TTS, and SER model utility better than VoicePrivacy Challenge baselines while improving privacy protection.
Details
Large-scale speech datasets are essential for training ASR, TTS, and SER systems — but they contain privacy-sensitive speaker identity and personal information. Existing anonymization approaches either strip acoustic features that degrade downstream utility, or fail to protect linguistic content (personally identifiable information in transcriptions). This paper proposes a two-stage framework. For content privacy, a generative speech editing model seamlessly replaces personally identifiable information (PII) in the audio without breaking acoustic continuity. For voice privacy, F3-VA (a flow-matching-based anonymization framework) produces diverse, distinct anonymized speakers using a three-stage design. Evaluation goes beyond the standard approach of testing anonymized speech with pretrained models. The authors train ASR, TTS, and SER models from scratch on anonymized data, then measure downstream task performance — a more realistic proxy for data utility under real training conditions. Privacy is assessed using both acoustic and content-based speaker verification metrics. Results show stronger privacy protection with minimal utility degradation compared to VoicePrivacy Challenge baselines. The framework is relevant for organizations that need to share or use speech datasets while complying with GDPR, HIPAA, or similar regulations.