
With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents
NVIDIA has put the NVIDIA Groq 3 LPX system into full production to extend the Vera Rubin NVL72 platform for agentic AI. The system utilizes extreme codesign across compute, networking, and inference to increase token generation speed.
Why it matters
Reducing latency in token generation allows AI agents to reason and respond faster, enabling more responsive real-time applications like autonomous coding systems for enterprises.
The details
- Groq 3 LPX achieved 3,400 output tokens per second using Gemma 4 31B.
- Spectrum-X Multiplane enables a flat network that scales to 512,000 GPUs.
- SpaceXAI is adopting NVIDIA Vera CPUs to accelerate orchestration and code execution.
Get the weekly recap
The stories like this one, picked and explained — once a week, straight to your inbox.