How XPUs Meet a World-Class AI Factory

| Source: NVIDIA Blog

Tags: NVIDIA, NVLink Fusion, XPU, AI infrastructure, custom silicon, hyperscalers, inference

NVIDIA's NVLink Fusion lets hyperscalers plug custom XPU silicon into NVIDIA's proven AI infrastructure stack, delivering 3x lower end-to-end latency and 10x higher packet rate vs. off-the-shelf Ethernet. The program targets cloud builders who want custom silicon advantages without rebuilding networking and software stacks from scratch.

Details

NVIDIA's NVLink Fusion program connects custom XPU accelerators — built by hyperscalers or AI-native companies — directly into NVIDIA's AI infrastructure rather than requiring teams to build scale-up networking and production software from scratch. The pitch is focused on time-to-market and risk reduction: innovate on the XPU, inherit the rest. Sixth-generation NVLink delivers end-to-end XPU-to-XPU latency 3x lower than alternative Ethernet solutions with 10x higher packet rate. Current 72-accelerator NVLink domains are on the roadmap to scale to 1,152 accelerators with co-packaged optics — relevant for hyperscalers running trillion-parameter models or mixture-of-experts architectures where scale-up bandwidth is a bottleneck. NVLink Fusion also includes NVLink-C2C, which connects XPUs to NVIDIA Vera CPUs or third-party CPUs at 6x the energy efficiency of PCIe — important for agentic systems where control and compute must stay tightly coupled with minimal overhead. This is an official NVIDIA blog, so treat performance claims as vendor-reported rather than independently benchmarked. The strategic angle is real: hyperscalers like Amazon (Trainium), Google (TPU), and Microsoft (Maia) want custom silicon but face significant integration costs. NVLink Fusion is NVIDIA's answer to keeping them in its ecosystem.