ACE: Adapter Consolidation across Experts for Parameter-Efficient Fine-Tuning of MoE LLMs

| Source: arXiv AI

Tags: MoE, LoRA, PEFT, fine-tuning, LLM, mixture-of-experts, EMNLP

ACE replaces per-expert LoRA adapters in mixture-of-experts LLMs with group-shared higher-rank modules, achieving the highest accuracy among parameter-matched PEFT methods on 3 of 4 MoE backbones while delivering 1.31–1.48× training speedup with no memory increase.

Details

Fine-tuning mixture-of-experts (MoE) LLMs with LoRA has a structural problem: attaching a separate adapter to each expert fragments capacity across many narrow low-rank updates, creates sparse and imbalanced gradient supervision under sparse routing, and decomposes execution into many small matrix multiplications (GEMMs) that are inefficient on modern hardware. ACE (Adapter Consolidation across Experts), accepted to EMNLP 2026, addresses this by identifying when expert-specific adapters become functionally similar during training — a sign of redundancy. It groups these redundant experts and replaces their individual adapters with a single shared higher-rank LoRA module at the same total parameter budget. The consolidated adapters are then executed as fewer, larger group-level GEMMs. Evaluated across 12 datasets and 4 MoE backbones, ACE achieves the highest mean accuracy among parameter-matched PEFT methods on 3 backbones, delivers 1.31× to 1.48× wall-clock training speedup over standard expert-wise LoRA, and does not increase peak memory usage. For teams fine-tuning MoE models like Mixtral or similar architectures, ACE appears to be a drop-in improvement over naive per-expert LoRA. Code is publicly available.