Building a Production Greek-English Speech Recognizer
| Source: arXiv AI
Tags: ASR, speech recognition, Greek, bilingual models, ROVER ensemble, production ML, Open ASR Leaderboard
A detailed engineering post-mortem documents building Sophea, a bilingual Greek-English ASR system: 23 training iterations, a data pipeline that recovered 88% of previously discarded Greek audio, and a 3-model ROVER ensemble that passed all 9 production quality gates with 4.26% average English WER.
Details
Sophea is a production ASR system for Greek and English, built over multiple months across 23 training iterations and two model architectures. The paper reports that no single training-data configuration passed all nine production gates simultaneously — English WER, Greek WER (noisy), language identification, and hallucination on non-speech were in tension throughout development. The most practically instructive contribution is the data pipeline. An audio-quality filter calibrated against in-domain anchors reduced the share of discarded Greek audio from 98.7% to 10.6% — recovering 88% of the dataset that a generic filter had thrown away. A pre-registered ablation isolated a hallucination defect to a single training-data package, demonstrating the value of controlled ablation design in production ML. The final model is a 3-model ROVER ensemble, which increased gate coverage from 4–7/9 (individual models) to 9/9 and reduced overlapping-speech WER from 53.35% to 37.87% (29% relative improvement). The ensemble model is listed as sophea/asr-k1 (preview) on the Open ASR Leaderboard with 4.26% average WER across eight English test sets and 25.88% on live Greek noisy-environment traffic. The authors document five cases where a measurement tool produced a plausible but wrong result — a rare candid account of evaluation tooling failures in production ASR development.