The cloud provider Nebius Group has signed on as the first customer to deploy the new processors within its production environment.
This technology addresses the critical challenge of decode latency, or the frustrating delays that occur when AI agents perform complex tasks like reasoning or coding.
By separating the massive data processing handled by graphics processing units (GPUs) from the high-speed generation of text or code, Nvidia claims the new accelerators can deliver record-breaking speeds of 3,400 tokens per second.
According to the company, this infrastructure allows multi-step agentic workflows that previously took hours to be completed in just minutes.
The Groq 3 LPX was developed using technology Nvidia licensed for $20 billion from the startup Groq Inc.
to focus specifically on inference—the process of running live AI models rather than training them.
Alongside this rollout, SpaceX has committed to using Nvidia’s Vera central processing units (CPUs) to manage complex simulations and tasks across both Earth-based data centers and orbital satellites.
Additionally, Nvidia introduced a new Ethernet architecture, Spectrum-X Multiplane, designed to link and scale massive clusters of up to 512,000 GPUs.