With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents
| Source: NVIDIA Blog
Tags: NVIDIA, Vera Rubin, Groq 3 LPX, inference, agentic AI, CoreWeave, SpaceXAI, AI infrastructure
NVIDIA's Groq 3 LPX is now in full production within the Vera Rubin NVL72 rack-scale system, hitting 3,400 output tokens/second at 4x the throughput of competing platforms for 100K-token agentic workloads. SpaceXAI, CoreWeave, and Nebius have committed to or deployed the platform.
Details
NVIDIA announced that Groq 3 LPX — its low-latency inference processor designed for agentic workloads — is now in full production as part of the Vera Rubin NVL72 rack-scale system. The headline benchmark: running Gemma 4 31B on Artificial Analysis, the system delivers 3,400 output tokens per second for 100,000-token long-context tasks at 4x the throughput of the nearest competing platform. The timing reflects a deliberate infrastructure pivot. As AI applications shift from training-heavy to inference-heavy agentic deployments, the bottleneck moves from peak GPU throughput to sustained token generation at long context windows. NVIDIA is explicitly positioning Vera Rubin + Groq 3 LPX for this emerging demand pattern, which is structurally different from what previous GPU generations were benchmarked against. Early adopters lend credibility. SpaceXAI announced NVIDIA Vera CPUs will power its next-generation agentic AI. CoreWeave has deployed Spectrum-X Multiplane into production — a multi-switch flat network topology connecting Vera Rubin racks for high-bandwidth, lossless AI networking. Nebius is the first cloud provider to adopt Groq 3 LPX commercially. Important naming note: Groq 3 LPX is NVIDIA's product — it is not affiliated with Groq Inc., the LPU inference startup. The shared name will cause confusion in the market.