Continuous-Time Acoustic Modelling with Neural Controlled Differential Equations

| Source: arXiv AI

Tags: TTS, speech synthesis, acoustic modeling, expressive speech, neural differential equations

Neural controlled differential equations (CDEs) applied to TTS duration modeling produce continuous-time hidden states whose values evolve with phonetic content, improving emotion intensity tracking in synthesized speech — accepted at IEEE SLT 2026.

Details

Text-to-speech systems conventionally expand phone encoder states to frame-level inputs using predicted durations — a discrete step that changes when states appear but not their values. This limits how well a model captures how phonetic content and timing jointly shape the acoustic output. This paper reframes duration modeling as a CDE control path: the phone representation parameterizes a continuous-time trajectory, and a neural acoustic vector field produces hidden states that evolve with phonetic content and duration-derived timing. The trajectory is sampled at discrete points for integration into a standard acoustic decoder. Evaluated on emotion-expressive TTS, CDE models with one phone per step improve rank-order agreement between synthesized and reference emotion intensity while maintaining comparable quality to a strong baseline. Half-phone step sizes reveal a tradeoff: finer temporal resolution improves style tracking but affects absolute calibration accuracy. The work is accepted at IEEE SLT 2026. It represents an elegant mathematical reformulation of TTS duration modeling with moderate empirical gains.