DriftSE: Speech Enhancement with Generative Drifting

| Source: arXiv AI

Tags: speech enhancement, generative models, audio AI, diffusion alternatives, latent diffusion

DriftSE proposes a one-step generative speech enhancement framework formulated as a latent distribution equilibrium problem, using a drifting field to align noisy speech distributions to clean speech — potentially faster than multi-step diffusion approaches.

Details

Speech enhancement — removing noise from degraded audio — has seen strong progress with diffusion models, but diffusion's multi-step inference is slow for real-time applications. DriftSE approaches the problem differently: it formulates enhancement as a latent distribution equilibrium problem and solves it in a single generative step. The 'drifting field' during training aligns the noisy speech latent distribution toward the clean speech distribution, rather than learning a denoising trajectory step by step. This framing enables one-step inference, removing the latency bottleneck of diffusion-based methods. The abstract was partially extracted, so specific benchmark results (DNSMOS, PESQ, STOI scores) and comparisons to diffusion baselines are not available from this feed entry. The approach is novel in its problem formulation as a distribution equilibrium problem rather than a denoising process. Authors are from Victoria University of Wellington and collaborators. Relevant to ASR preprocessing, hearing aid applications, and communications systems teams building on top of speech enhancement.