Traceable TTS: Toward Watermark-Free TTS with Strong Traceability

| Source: arXiv AI

Tags: TTS, speech-synthesis, model-attribution, deepfake-detection, audio-AI, watermarking

A watermark-free TTS traceability framework uses joint training with a discriminator to attribute synthetic speech to its source model without embedding watermarks — the first approach to achieve strong traceability while preserving (and slightly improving) audio quality.

Details

Synthetic speech can now mimic human voices convincingly enough to enable fraud, deepfakes, and voice cloning at scale. Tracing synthetic speech back to its source model is a key defense — but existing methods either embed explicit watermarks that degrade quality or rely on vocoder-level markers that are fragile to common processing. This paper proposes a watermark-free alternative: instead of marking the output, train the TTS model itself alongside a discriminator using a joint training method that improves traceability generalization while preserving audio quality. The discriminator effectively shapes the model's learned representation to make attribution possible without injecting external signals. The authors claim this is the first watermark-free TTS system with strong traceability. Audio quality is preserved and slightly improved versus baseline. Code will be released upon paper acceptance. The approach matters because existing watermarking methods can be stripped by audio processing (resampling, compression, noise addition). A system where traceability is baked into the model's generative process is harder to defeat — but also raises questions about whether this would actually hold against adversarial processing at inference time, which the paper does not fully address. The submitted-July-2025 timestamp suggests this is earlier-stage work that has since been developed further.