X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System

| Source: arXiv AI

Tags: speech-to-speech translation, real-time, speaker diarization, TTS, streaming ASR, open-source, multilingual

X-Translator is an open-source real-time speech-to-speech translation system that preserves speaker voice identity across languages—using streaming ASR, machine translation, and a session-level speaker prompt manager with online speaker diarization, evaluated against proprietary APIs on OpenSTBench.

Details

Real-time speech-to-speech translation (S2ST) faces competing constraints: translation quality, latency, speech naturalness, and speaker identity preservation. While proprietary APIs increasingly offer S2ST, few open systems document how they handle multi-speaker conversations over extended sessions.\n\nX-Translator takes a modular cascaded approach: streaming ASR converts speech to unstable partial hypotheses, an incremental segment commitment module converts those into translation-ready units, machine translation processes them, and a prompt-conditioned TTS synthesizes the target speech. A session-level speaker prompt manager binds source speech spans to speaker-specific voice prompts, preserving distinct voices throughout the conversation.\n\nThe system is evaluated with OpenSTBench across translation quality, speech naturalness, latency, long-form voice stability, speaker preservation in multi-speaker settings, and multilingual translation quality. Proprietary S2ST APIs serve as behavioral baselines. Code and demo are publicly available.\n\nFor developers building real-time translation applications, X-Translator provides an open, reproducible baseline for understanding the deployment trade-offs that proprietary systems obscure. The speaker prompt manager architecture is particularly relevant for applications like conference interpretation or video dubbing where voice consistency matters.