Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents

| Source: NVIDIA Blog

Tags: NVIDIA, Vera Rubin, Blackwell, agentic AI, AI inference, GPU efficiency, SemiAnalysis, DeepSeek

On-silicon benchmarks show NVIDIA Vera Rubin NVL72 delivers 30x higher throughput per megawatt and 35x lower token costs vs. GB300 NVL72 on agentic workloads — measured using SemiAnalysis AgentX, which replays real recorded coding agent sessions rather than synthetic inference runs.

Details

NVIDIA published on-silicon performance data showing Vera Rubin NVL72 achieves 30x higher throughput per megawatt and 35x lower token cost compared to GB300 NVL72 on agentic workloads. The benchmark methodology matters: NVIDIA used SemiAnalysis AgentX, which replays real recorded agentic coding trajectories — preserving actual context growth, tool calls, and sub-agent spawning patterns rather than simplified synthetic tests. The efficiency gains are driven by how agentic workloads differ structurally from chat or summarization. OpenRouter data cited in the post shows agentic tasks consume 15x more tokens than simple chat requests, with context that accumulates across steps and can reach hundreds of thousands of tokens. Traditional GPU benchmarks optimized for short-sequence throughput don't capture this well. For context on the generational progression: GB300 NVL72 already delivers 15x better throughput per megawatt than Hopper architecture on DeepSeek V4 Pro. Vera Rubin's 30x figure is measured against GB300 NVL72, not Hopper — the cumulative efficiency jump over older hardware is substantial for infrastructure planning. The 35x lower token cost figure is directly actionable for enterprises budgeting agentic deployments. For power-constrained data centers, 30x more work per megawatt means the same power envelope supports a dramatically larger agentic workload at equivalent cost.