Bridging the Modality Gap in Long-Form Clinical Audio: A Comparative Study of Lightweight and Heavyweight End-to-End SOAP Generation
| Source: arXiv AI
Tags: clinical NLP, medical documentation, SOAP notes, speech-to-text, healthcare AI, multimodal
The ASLP team's fully end-to-end system generates structured SOAP clinical notes directly from long doctor-patient audio conversations, consistently outperforming cascaded ASR+LLM baselines at both 3B and 30B parameter scales using a 1.41-million-sample training corpus.
Details
Automating clinical documentation is one of the highest-value AI applications in healthcare, but end-to-end audio-to-text models have historically underperformed cascaded pipelines that first transcribe with ASR then summarize with LLMs. This paper presents a full E2E system for the BeTraC 2026 challenge that processes raw audio directly into SOAP notes without intermediate transcripts. The team built a 1.41-million-sample multi-task corpus and used a three-stage pipeline: domain pre-training, supervised fine-tuning, and reward optimization. Testing at Lightweight (3B) and Heavyweight (30B) scales, the 30B version substantially improves medical concept extraction and summarization quality. The E2E approach consistently outperforms representative cascaded baselines, validating that direct multimodal optimization reduces information loss and hallucination risk that intermediate transcription introduces. Each training stage progressively enhances performance, with reward optimization providing the final alignment gains. This is relevant for clinical documentation platforms competing to reduce physician EHR burden. The challenge context limits direct generalizability claims, but the architecture choices and corpus scale are informative for teams building production systems.