NVIDIA says its Vera Rubin NVL72 platform can deliver up to 30 times higher throughput per megawatt than Grace Blackwell NVL72 in demanding agentic AI workloads, while also reducing the cost per million tokens by as much as 35 times.
The results come from NVIDIA’s latest on silicon testing using the SemiAnalysis AgentX benchmark, which measures real world agentic coding inference across models including Kimi K3, MiniMax M3, GLM5.3, Qwen3.5 and DeepSeek V4 Pro.
The benchmark focuses on more than raw token generation. It also measures end to end responsiveness, time to first token and how much useful output a platform can deliver within a fixed power budget.
Vera Rubin targets much higher efficiency
| Metric | Reported result |
|---|---|
| Vera Rubin throughput gain vs Blackwell | Up to 30x per MW |
| Vera Rubin token cost reduction | Up to 35x |
| Vera Rubin interactivity | Up to about 280 tokens per second per user |
| Blackwell peak interactivity | Below 180 tokens per second per user |
| Blackwell throughput gain vs Hopper | Up to 15x in one DeepSeek workload |
| Blackwell token cost reduction vs Hopper | Up to 10x |
| Blackwell gain in Kimi K3 workload | Up to 80x throughput per MW |
| DSX MaxLPS power benefit | Up to 40% more GPUs per MW budget |
NVIDIA says Vera Rubin reached the 30x throughput improvement in DeepSeek V4 Pro 1.6T testing.
At around 160 tokens per second per user, Vera Rubin reportedly maintained substantially higher throughput per megawatt than Grace Blackwell NVL72.
The platform also reached roughly 280 tokens per second per user in the benchmark, while Blackwell remained below 180.
Blackwell already shows large gains over Hopper
Before comparing Vera Rubin with Blackwell, NVIDIA also highlighted how much its current Blackwell platform has improved over Hopper.
The GB300 NVL72 reportedly delivered up to 15 times higher throughput per megawatt than an H200 NVL8 system in DeepSeek V4 Pro 1.6T.

NVIDIA also claims Blackwell can reduce cost per million tokens by as much as 10 times.
In a Kimi K3 2.8T workload, the difference was even larger, with Blackwell reaching up to 80 times the throughput per megawatt of Hopper while sustaining around 215 tokens per second per user.
These are NVIDIA supplied benchmark results, so independent testing will still be important for understanding how the platforms compare across broader workloads.
Agentic AI changes how infrastructure is measured
Traditional inference benchmarks often focus on isolated throughput or latency.
Agentic AI workloads are more complex because agents can issue repeated long context requests, call tools, generate large amounts of output and interact with external systems.
AgentX therefore tracks several metrics, including end to end interactivity, standard interactivity, total request latency and time to first token.
That matters because very high throughput is less useful if each request takes too long to start or complete.
NVIDIA is positioning Vera Rubin specifically around this type of workload, where power efficiency and continuous agent activity become more important than peak performance alone.
Token costs could fall sharply
The 35x lower token cost claim is one of the most important parts of the Vera Rubin results.
Lower cost per million tokens could allow operators to run more AI agents continuously without increasing infrastructure spending at the same rate.
NVIDIA also points to its DSX MaxLPS power management technology, which coordinates power use across GPUs, racks and workloads.
The company says this can allow AI factories to deploy up to 40% more GPUs within the same megawatt power envelope.
The benchmark does not include the full Vera Rubin platform
NVIDIA notes that these results do not yet include the full contribution of Vera CPUs for tool calling.
The testing focuses mainly on the Vera Rubin compute platform rather than the complete seven chip architecture NVIDIA describes as its full stack AI factory design.
That broader system also includes Groq 3 LPX inference hardware, Vera CPUs, BlueField 4 storage infrastructure and Spectrum 6 networking.
NVIDIA says production has now started across these major components.
If the company’s benchmark claims hold up in independent testing, Vera Rubin could represent a major jump in efficiency for agentic AI, with much higher throughput per watt and far lower token costs than Blackwell.



Discussion (0)
Be the first to comment.