Earnings25: A Comprehensive 500-Hour Speech Benchmark for Finance

| Source: arXiv AI

Tags: ASR, speech recognition, finance, benchmark, Whisper, Parakeet

Earnings25 releases a 498-hour ASR benchmark built from S&P 500 Q4 2025 earnings calls, including speaker role and industry metadata that enable evaluation beyond standard word error rate — with baselines for Whisper and Parakeet-TDT.

Details

ASR models are widely deployed in finance for transcribing earnings calls, investor meetings, and regulatory filings — but until now, no rigorous public benchmark covered real-world S&P 500 calls at scale. Earnings25 fills this gap with two test sets: a full 498-hour set from English-language S&P 500 Q4 2025 earnings calls, and a 46-hour industry-balanced set of 290 segments sampled from broader U.S. earnings calls in 2025. The benchmark includes structured metadata covering speaker roles (executives, analysts), industry labels, and call structure — enabling evaluation that goes beyond aggregate word error rate to assess speaker-specific and sector-specific performance. Baseline results are provided for Whisper and Parakeet-TDT using standardized scoring, giving a comparable starting point for teams evaluating ASR systems for financial applications. The study is concise (5 pages) and intended as a community resource rather than a modeling contribution. For teams building speech-to-text pipelines in finance, Earnings25 provides a practical evaluation foundation using real S&P 500 calls with reproducible baselines — something the field has lacked.