X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation
| Source: arXiv AI
Tags: speech LLM, model compression, distillation, Qwen3-ASR, LoRA, ASR, audio encoder
X-AuT compresses speech LLM audio encoders by progressively pruning layers and restoring quality via cross-scale distillation — reducing Qwen3-ASR audio-tower parameters by 20.7% while actually lowering macro-average error from 5.61% to 5.27% on a 16-layer model.
Details
Speech LLMs like Qwen3-ASR pair a large audio encoder with a language model backbone, and the audio encoder contributes meaningfully to inference cost. X-AuT targets this encoder specifically, using short behavioral probes to select which layer combinations to prune, then restoring performance with cross-scale distillation, LoRA adapters, and a transcript-consistency pipeline for training data quality. On Qwen3-ASR-0.6B, compressing from 18 to 16 audio-encoder layers actually reduces macro-average error from 5.61% to 5.27% on ten public Chinese-English benchmarks — an improvement, not just graceful degradation. The 14-layer model reaches 5.75% with 20.7% fewer audio-tower parameters, a reasonable trade-off for latency-sensitive deployment. Progressive pruning (18→14 layers) consistently outperforms direct pruning to 14 layers (5.75% vs. 6.73%), showing the gradual approach matters. The language model backbone stays frozen throughout, which makes this approach portable to other speech LLMs using similar architecture patterns. Accuracy effects vary meaningfully across benchmarks, suggesting practitioners need to evaluate on their target distribution before committing to a compressed configuration.