TokenMapper: A Step Toward Interoperable Speech Token Translation
| Source: arXiv AI
Tags: speech tokenization, voice AI, audio codecs, speech-to-speech, latency optimization, GLM-4-Voice, DualCodec
TokenMapper enables direct token-to-token translation between structurally different speech tokenizers (single vs. multi-codebook), reducing end-to-end latency by 4.8-94.5% versus waveform bridging — accepted at AACL-IJCNLP 2026.
Details
Modern voice AI systems chain together multiple speech models (ASR, LLMs, TTS), but each model uses its own audio tokenizer with a different vocabulary and codebook structure. Currently, transferring speech between models requires decoding to raw waveform audio and re-encoding — introducing latency (potentially hundreds of milliseconds) and potential quality loss.\n\nTokenMapper is a direction-aware framework for direct token-to-token translation in the discrete domain, handling structurally mismatched token spaces including single-to-multi-codebook and multi-to-single-codebook mappings. The key constraint: mappings must share an effective token rate, which maintains temporal alignment across models.\n\nExperiments on three tokenizer pairs (GLM-4-Voice, MiMi, DualCodec) show: translation WER approaches native reconstructions within 2.5-6.8% absolute WER, human MOS for TokenMapper outputs ranges from 2.29 to 4.39 (following the same trend as UTMOS automated scores), and end-to-end latency is reduced by 4.8-94.5% relative to waveform bridging — up to 972ms saved per utterance in some configurations.\n\nAccepted at AACL-IJCNLP 2026. For teams building conversational voice AI, speech-to-speech translation, or streaming voice agents, direct tokenizer interoperability is a meaningful engineering improvement.