A survey of AI-generated voices and their detection

| Source: arXiv AI

Tags: voice cloning, deepfake audio, TTS, fraud detection, audio security, synthetic speech

A comprehensive survey of AI voice generation and detection covers TTS and voice cloning through deepfake audio detection, highlighting recent voice cloning scams targeting businesses and political leaders as evidence that synthetic voice fraud has moved beyond theoretical concern.

Details

The survey by Sun, Yang, and Lyu covers the full arc of AI-generated speech: text-to-speech systems, voice cloning, and the detection methods designed to identify synthetic audio. It spans both technical foundations and the current state of the art, identifying open challenges, benchmark resources, and future research directions. The timing reflects a real escalation. The authors cite recent voice cloning scams targeting businesses and political leaders as evidence that the problem has become operationally serious—these attacks exploit the gap between how convincingly AI generates voice and how effectively detection systems can identify it. Unlike image and video deepfakes, voice fakes pose distinct challenges due to the complexity of phonetics, prosody, and auditory perception. The survey is positioned as a reference for future researchers, cataloging detection benchmarks and summarizing open problems. It does not propose a novel detection method, but offers a structured overview of what approaches exist, their limitations, and where significant gaps remain. For practitioners building voice-authentication or anti-fraud systems, the paper provides a practical map of the detection landscape—including which benchmarks exist and what open challenges remain unresolved.