Speech to Text News and AI Updates
Follow Speech to Text developments across AI companies, labs, and open-source projects.
Latest Speech to Text news, research, benchmarks, product releases, and industry adoption updates.
Latest Articles
- Bridging the Modality Gap in Long-Form Clinical Audio: A Comparative Study of Lightweight and Heavyweight End-to-End SOAP Generation — The ASLP team's fully end-to-end system generates structured SOAP clinical notes directly from long doctor-patient audio conversations, consistently outperforming cascaded ASR+LLM baselines at both 3B and 30B parameter scales using a 1.41-million-sample training corpus.
- Building a Production Greek-English Speech Recognizer — A detailed engineering post-mortem documents building Sophea, a bilingual Greek-English ASR system: 23 training iterations, a data pipeline that recovered 88% of previously discarded Greek audio, and a 3-model ROVER ensemble that passed all 9 production quality gates with 4.26% average English WER.
- TokenMapper: A Step Toward Interoperable Speech Token Translation — TokenMapper enables direct token-to-token translation between structurally different speech tokenizers (single vs. multi-codebook), reducing end-to-end latency by 4.8-94.5% versus waveform bridging — accepted at AACL-IJCNLP 2026.
- Not All Speech Is Intent: Adaptive Self-Correcting Inference Layer for Post-ASR False Wake-Up — ASCIL is a post-ASR correction framework that cuts false wake-up errors by 54.27% relative on a session-disjoint subset, using acoustic embeddings, hesitation/silence signals, and personalized learning from past misclassifications — with under 60ms added latency.
- DriftSE: Speech Enhancement with Generative Drifting — DriftSE proposes a one-step generative speech enhancement framework formulated as a latent distribution equilibrium problem, using a drifting field to align noisy speech distributions to clean speech — potentially faster than multi-step diffusion approaches.
- Beyond Prompting: Efficient and Robust Contextual Biasing for Speech LLMs via Logit-Space Integration (LOGIC) — LOGIC injects entity probability boosts directly into the decoding layer of speech LLMs — achieving 9% relative reduction in Entity Word Error Rate across 11 multilingual locales with Phi-4-MM, while maintaining constant-time complexity regardless of entity list size.
- Beyond Word Error Rate: A Switch Aware Evaluation of ASR and Audio Language Models on English Yoruba Code-Switched Speech — A rigorous evaluation of 11 ASR and audio LM systems on English-Yoruba code-switched speech finds that aggregate WER masks critical failures: the top-WER ASR model is statistically indistinguishable from an audio LM on WER but far worse on every switch-localized metric.
- Towards AI-Driven Policing: Interdisciplinary Knowledge Discovery from Police Body-Worn Camera Footage — Researchers propose a multimodal AI framework — combining video, audio, and NLP — to analyze Rochester Police Department body-worn camera footage, automatically classifying behavioral dynamics like de-escalation and disrespect to support law enforcement accountability reviews.
- X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation — X-AuT compresses speech LLM audio encoders by progressively pruning layers and restoring quality via cross-scale distillation — reducing Qwen3-ASR audio-tower parameters by 20.7% while actually lowering macro-average error from 5.61% to 5.27% on a 16-layer model.
- OpenAI's GPT-Live-1 API lets developers build apps that talk and listen at the same time — OpenAI's GPT-Live-1 API brings full-duplex speech — simultaneous listening and speaking — to developers at $0.05/minute. Compared to its predecessor, interactivity scores jump from 45.4% to 80.1%, turn-taking latency drops from 1.4s to 0.8s, and tool-calling accuracy rises from 60% to 87%.