OpWeave: Flexible Operator Disaggregation for Heterogeneous LLM Serving

| Source: arXiv AI

Tags: LLM serving, inference optimization, GPU clusters, disaggregated serving, vLLM, heterogeneous hardware

OpWeave cuts LLM serving costs by up to 1.89x on heterogeneous GPU clusters by disaggregating operators (attention vs FFN/MoE) across device groups — with an analytical cost model to determine when disaggregation actually helps vs. hurts.

Details

LLM serving has moved toward disaggregating prefill and decode, but recent research goes further: separating attention from FFN or MoE computation across different hardware. The challenge is that existing operator-disaggregated serving (ODS) systems fix operator boundaries and lack a principled way to determine when disaggregation actually reduces cost versus when colocation is cheaper.\n\nOpWeave presents an end-to-end ODS framework built on three components: an analytical cost model that bounds gains of both homogeneous and heterogeneous ODS over colocation; a regularity-aware planner that jointly optimizes operator partitioning and deployment configuration (keeping search tractable for hybrid-attention models); and a vLLM-based runtime that executes synthesized plans across heterogeneous device groups.\n\nEvaluation shows up to 1.78x cost reduction on homogeneous GPU clusters and 1.89x on heterogeneous clusters relative to the best feasible baseline, while meeting latency SLOs. The cost model is the key contribution: it tells you in advance whether disaggregation is worth the added infrastructure complexity for a given model and hardware mix.\n\nDirectly relevant to infrastructure teams running large LLM serving fleets on multi-GPU or mixed-hardware clusters.