NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory
| Source: NVIDIA Blog
Tags: NVIDIA, NVLink Fusion, NVHBM, HBM, Trainium, Amazon, Annapurna Labs, AI infrastructure
NVIDIA expanded NVLink Fusion with NVHBM, a custom high-bandwidth memory that moves the controller into the HBM stack rather than the XPU die, delivering 30% more bandwidth, 15% lower power, and 25% more compute die area. Amazon's Annapurna Labs is first to adopt it for Trainium4.
Details
NVIDIA announced NVHBM, a next-generation high-bandwidth memory technology that rethinks HBM architecture at the die level. Instead of placing the memory controller on the XPU compute die, NVHBM integrates NVIDIA's custom controller into the HBM base die itself, built on the same technology NVIDIA uses for its own future GPUs. The result: up to 30% greater memory bandwidth, 15% lower HBM power consumption, and 25% more compute die area freed up on the XPU for actual AI compute. NVHBM extends the NVLink Fusion platform, which lets hyperscalers build semi-custom AI accelerators (XPUs) that connect to NVIDIA's rack-scale infrastructure including NVLink chiplets, switches, and MGX systems. NVIDIA is standardizing NVHBM across multiple memory vendors, cutting qualification overhead and accelerating time-to-market for custom chips. Amazon's Annapurna Labs is the launch partner. Trainium4 will support NVLink Fusion and NVHBM, meaning Amazon's custom AI chips and NVIDIA GPUs will share a common rack-scale architecture at AWS. This deepens AWS's hybrid silicon strategy rather than a fully proprietary stack. For enterprises, NVHBM's bandwidth gains are directly relevant to memory-bound inference workloads. The NVLink Fusion model positions NVIDIA as indispensable networking and memory infrastructure even for hyperscalers building their own compute dies, a deliberate strategy to maintain platform lock-in through the connectivity and memory layer.