NVIDIA has moved its Groq 3 LPX inference accelerator into full production, adding a dedicated low latency decode engine to its broader Vera Rubin AI platform.
The company says Groq 3 LPX is designed to accelerate token generation for agentic AI workloads, where a single task may involve hundreds or thousands of inference steps and fast response times can be as important as total throughput.
In one demonstration, NVIDIA paired Vera Rubin NVL72 with Groq 3 LPX and reported 3,400 tokens per second while running the Gemma 4 31B reasoning model with a 100,000 token context window.
According to the supplied results, this was the fastest recorded performance for that model at the time of testing.
Groq 3 LPX focuses on latency sensitive token generation
| Area | Detail |
|---|---|
| Product | NVIDIA Groq 3 LPX |
| Status | Full production |
| Main role | AI inference and token generation |
| Platform | Vera Rubin |
| Demonstrated model | Gemma 4 31B |
| Reported speed | 3,400 tokens per second |
| Context window | 100,000 tokens |
| Workload focus | Agentic AI and reasoning |
| Cloud deployment | Nebius Token Factory planned |
NVIDIA is using Groq 3 LPX as a complementary accelerator rather than a replacement for Rubin GPUs.
Rubin GPUs handle large context processing and other compute intensive parts of inference, while LPX focuses on the decode stage where generated tokens need to be produced quickly and consistently.
That division of work is intended to improve responsiveness without forcing every stage of inference onto the same hardware.
Vera Rubin and LPX split inference workloads
Large language model inference generally contains different phases with different hardware requirements.
The prefill stage processes the incoming prompt and can require significant parallel compute, especially with long context windows.

The decode stage then generates tokens sequentially, making latency more important.
NVIDIA’s approach is to let Rubin GPUs handle the broader context processing while Groq 3 LPX accelerates the latency sensitive token generation stage.
This is particularly relevant for AI agents, which may repeatedly generate code, call tools, inspect results and begin another inference cycle.
Reducing the time spent during each decode phase can shorten the overall completion time of those multi step tasks.
NVIDIA claims four times faster response in some workloads
The company also says Groq 3 LPX can provide up to four times faster response performance than the nearest alternative platform in selected agentic AI tests.
NVIDIA argues that this can reduce coding tasks from hours to minutes in some scenarios.
Those figures are vendor supplied results and depend heavily on the model, workload, serving configuration and comparison point, so they should not be treated as universal performance expectations.
Independent testing will be important for determining how LPX performs across a wider range of models and real production environments.
Full production expands NVIDIA’s Vera Rubin rollout
Groq 3 LPX entering full production follows NVIDIA’s production ramp for other major Vera Rubin components.
The company has also moved Vera CPUs and Vera Rubin server systems into production, indicating that the wider platform is transitioning from development and demonstration into deployment.
Vera Rubin is being positioned as a modular AI factory platform rather than a single accelerator.
Different hardware can be assigned to different parts of an AI workload, including CPU orchestration, large scale GPU compute, networking and high speed inference generation.
Groq 3 LPX adds another specialized layer to that design.
Nebius plans to use Groq 3 LPX
Cloud infrastructure provider Nebius plans to deploy Groq 3 LPX through its Nebius Token Factory.
The goal is to expose the faster generation hardware through the same API environment developers already use, avoiding the need to rewrite applications around a separate software stack.
That type of integration could be important if NVIDIA wants LPX to gain adoption beyond its own reference systems.
Enterprises and AI cloud providers generally need new accelerators to work within existing serving frameworks and operational tools rather than requiring a completely separate deployment model.
Agentic AI is driving demand for specialized inference hardware
The broader significance of Groq 3 LPX is NVIDIA’s increasing focus on separating inference into specialized stages.
As agentic AI workloads become more complex, total model performance alone may be less important than how quickly infrastructure can cycle through repeated inference steps.
A system that processes long prompts efficiently but generates output slowly can still make an agent feel unresponsive.
By pairing Rubin GPUs with LPX hardware, NVIDIA is trying to optimize both sides of that problem.
With Groq 3 LPX now in full production, the company is moving toward a Vera Rubin architecture where general AI compute and ultra fast token generation can be handled by different accelerators inside the same infrastructure.



Discussion (0)
Be the first to comment.