The results follow Nvidia's late 2024 acquisition of Groq, a startup specializing in Language Processing Units (LPUs) that utilize high-speed static random-access memory (SRAM) to overcome the data transfer bottlenecks common in traditional AI chips.
This speed increase is significant for the development of "agentic" AI—autonomous tools or assistants that perform complex reasoning and multi-step tasks.
High-speed inference (the process of a model generating a response) allows these agents to process more information and take more actions within a standard time frame.
To achieve this, Nvidia is integrating these LPUs into its broader Vera Rubin AI infrastructure platform.
The company’s strategy involves a heterogeneous architecture where standard graphics processing units (GPUs) handle the initial "prefill" stage of a prompt, while the Groq 3 LPUs manage the memory-intensive "decode" phase of generating text.
While the benchmark demonstrates high performance for a 31-billion parameter model, the architecture faces scaling challenges due to the physical limits of SRAM.
Because each Groq 3 LPU contains only 500 MB of memory—significantly less than the 288 GB found on Nvidia’s Rubin GPUs—models must be distributed across dozens of chips using high-speed Ethernet.
For example, running the Gemma 4 31B model requires at least 64 LPUs.
Despite these hardware requirements, the technology is moving into active use, with the European cloud provider Nebius recently becoming one of the first customers to deploy the combined GPU and LPU systems in its data centers.