SeqMoE: Toward Full-Load Performance via Predictive and Graph-Compatible MoE Offloading

| Source: arXiv AI

Tags: MoE, inference optimization, model offloading, Mixture-of-Experts, GPU memory, CUDA, DeepSeek

SeqMoE achieves 96.97% expert hit rate and 80.22% of full-load inference performance while keeping only 45% of MoE expert weights in device memory, using sequence-to-sequence prediction for expert prefetching — making large Mixture-of-Experts models practical on memory-constrained hardware.

Details

Mixture-of-Experts models have a structural advantage for offloading: only active experts need GPU memory. But translating that into near-full-load performance requires predicting which experts will be needed far enough in advance to prefetch them without stalling computation. Prior work has not fully cracked this. SeqMoE combines four innovations. First, it recasts expert activation prediction as sequence modeling — forecasting which experts will be needed multiple steps ahead across multiple layers, providing a longer and more reliable prefetch window than single-step prediction. Second, it formulates prefetch scheduling as Job Sequencing with Deadlines to maximize expected expert hits given bandwidth constraints. Third, it uses a probabilistic Belady eviction policy that leverages the recursive nature of sequence modeling for future-aware cache management. Fourth, it develops a graph-compatible offloading runtime that enables end-to-end CUDA graph capture without blocking compute. The results: with 45% expert residency (55% of weights offloaded), SeqMoE averages a 96.97% expert hit rate and achieves 80.22% of full-load performance — a substantial improvement over prior MoE offloading approaches. For practitioners running large MoE models (like Mixtral, DeepSeek-MoE, or similar) on memory-constrained hardware, this is directly useful.