Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts

| Source: arXiv AI

Tags: Mixture-of-Experts, MoE, inference efficiency, EMNLP 2026, LLM optimization, sparse models

EMNLP 2026 paper introduces Layer-wise Distribution Alignment (LDA), a calibration-based inference-time fix for Sparse MoE models that recovers most accuracy lost when using fewer active experts—no retraining required, negligible overhead.

Details

Sparse Mixture-of-Experts (SMoE) models like Mixtral and similar LLMs use a fixed top-k expert selection per token. Reducing k at inference saves compute without retraining, but prior methods see accuracy degrade. Kim, Soro, and Hwang (EMNLP 2026) identify two separate causes: loss of expert capacity (expected) and a distribution mismatch caused by the reduced routing increasing RMS scale and variance of layer outputs (overlooked by prior work). Their fix, Layer-wise Distribution Alignment (LDA), applies lightweight layer-wise calibration statistics computed offline to align reduced-routing representations with the full-routing distribution. It is a pure inference-time intervention requiring no model retraining and adding negligible computational overhead. Results span multiple SMoE LLMs, benchmarks, and routing strategies. LDA recovers much of the performance lost under reduced routing while preserving sparse-inference efficiency. The paper was accepted to Findings of EMNLP 2026. Practically, this means teams running MoE models can activate fewer experts per token during serving—saving FLOPs proportionally—with LDA as a calibration correction to keep accuracy close to the full-routing baseline.