mKernel: Fast Multi-GPU, Multi-Node Fused Kernels
| Source: arXiv AI
Tags: distributed training, multi-GPU, NVLink, RDMA, inference optimization, tensor parallelism, Ring Attention
mKernel achieves up to 1.88x speedup on Ring Attention across 16-GPU H200 clusters by fusing computation with NVLink and RDMA communication at tile granularity — targeting the communication bottleneck that limits distributed LLM training at scale.
Details
Communication overhead has become the dominant bottleneck in large-scale distributed training and inference. Existing approaches overlap computation and communication on separate CUDA streams, but this only partially hides the cost. mKernel takes a tighter approach: it fuses compute with NVLink (intra-node) and RDMA (inter-node) at tile granularity, transmitting each output tile as soon as it is produced. The key innovation is an on-GPU controller that adaptively partitions streaming multiprocessors (SMs) between compute and communication roles at runtime, since the optimal partition varies per kernel and input shape. Data traversal is structured hierarchically to minimize inter-node traffic, and network operations run directly from the GPU via a lightweight RDMA command queue — making it portable across InfiniBand and AWS EFA. The authors implement five kernels covering tensor parallelism, sequence parallelism (Ring Attention), and expert parallelism. On two 16-GPU H200 clusters: 1.72x speedup on GEMM+AllReduce, 1.88x on Ring Attention. A notable finding: GPUDirect Async (IBGDA) offered little benefit over host-assisted GPU-initiated communication, contrary to common assumptions. The library comes from UC Berkeley's RISE Lab (Ion Stoica, Scott Shenker), which has a track record of producing adopted distributed systems infrastructure.