Speech to Text News and AI Updates
Follow Speech to Text developments across AI companies, labs, and open-source projects.
Latest Speech to Text news, research, benchmarks, product releases, and industry adoption updates.
Latest Articles
- Auditing Self-Evolution in Financial Agents: Capability Gains, Security Drift, and Execution-Interface Mismatch — Auditing three self-evolving agent frameworks in simulated e-banking finds that capability gains reliably come with security drift: SkillOpt raises benign utility from 0.741 to 0.837 but unauthorized financial state changes climb to 0.685.
- Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets — Fair-ASR is a new jailbreak evaluation protocol that uses target API calls as the comparison axis — and when applied to 11 attacks, simple stochastic methods outperform complex LLM-based attacks under equal budget. ReCode, built from this analysis, achieves 85% attack success on GPT-5 with just 20 target calls.
- Language Family Matters: Evaluating LLM-Based ASR Across Linguistic Boundaries — Training one ASR connector per language family rather than per language reduces parameter count while improving cross-domain generalization in LLM-based speech recognition — validated on two multilingual LLMs and two speech corpora, accepted at EACL 2026.
- JailbreakSkill: Scaling Automated Red-Teaming with Reusable and Ever-Evolving Skills — JailbreakSkill packages jailbreak attacks into modular, evolving skills — lifting attack success rates by 17.5 pp on AdvBench and achieving a 48.6-point gain against GPT-5.4, with novel attack strategies (like reframing harmful requests as document completion) emerging automatically from the feedback loop.
- Lost in Phonation: Voice Quality Variation as an Evaluation Dimension for Speech Foundation Models — VQ-Bench reveals that leading speech foundation models systematically shift responses based on voice quality — attributing different agency, empathy, and leadership to modal vs. breathy vs. creaky voices — with gender asymmetries in salary and leadership endorsements; accepted at Interspeech 2026.
- Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition — A controlled benchmark of six ASR models on 165 hours of Nepali finds Whisper-Large-v3-Turbo and IndicWav2Vec tie at the top (14.76% vs 14.89% WER) despite a 9x parameter gap — while CTC decoders run up to 29x faster at equivalent accuracy, flipping the deployment preference for any latency-constrained application.
- Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed — OpenAI is previewing Ultrafast, an API tier running GPT-5.6 Sol at up to 750 output tokens per second — 14× standard speed — powered by Cerebras inference hardware, targeting latency-critical applications where throughput matters more than cost.
- Smart rings are looking like my kind of AI gadget — Sandbar's upcoming Stream ring — a minimal microphone-battery-Bluetooth wearable — offers a discreet alternative to smart glasses for AI voice interaction, with press-to-speak activation and a whisper mode for private, on-the-go AI queries.
- Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages — Indic DiarBench releases 108 hours of human-annotated multi-speaker audio across all 22 scheduled languages of India, establishing the first unified benchmark for joint speaker diarization and ASR in Indian languages — accepted at Interspeech 2026.
- Earnings25: A Comprehensive 500-Hour Speech Benchmark for Finance — Earnings25 releases a 498-hour ASR benchmark built from S&P 500 Q4 2025 earnings calls, including speaker role and industry metadata that enable evaluation beyond standard word error rate — with baselines for Whisper and Parakeet-TDT.