NVIDIA has shared the first performance figures for its Vera Rubin NVL72 platform, claiming a tenfold increase in token throughput compared with the current GB200 Grace Blackwell system at the same power level.
In an early DeepSeek R1 test, the Vera Rubin rack reportedly sustained 800,000 tokens per second while operating within a 150MW deployment. A GB200 NVL72 configuration delivered 80,000 tokens per second under the same stated power limit.
The result highlights NVIDIA’s growing focus on performance per watt rather than raw compute alone. For companies running large AI services, higher token throughput from the same power budget can reduce infrastructure costs and allow more customer requests to be processed without expanding power capacity.
These figures come from an early platform deployment and should not be treated as a universal performance comparison. Results will vary according to model size, precision, networking, software, batch configuration, and the type of AI workload being used.
Vera Rubin combines compute, networking, and storage at rack scale
Vera Rubin is not simply a new GPU. It is a complete hardware and software platform built from several specialised trays and interconnect technologies.
The NVL72 compute system includes Rubin accelerators, Vera processors, Grace BlueField components, and ConnectX 9 networking. Separate trays handle NVLink switching, storage, Ethernet networking, and specialised inference tasks.
| Vera Rubin component | Main role |
|---|---|
| Vera Rubin NVL72 compute tray | AI training and inference |
| NVLink 6 switch tray | High speed communication within the rack |
| Vera CPU tray | General processing and system coordination |
| Spectrum 6 switch tray | Scale out Ethernet networking |
| Groq 3 LPX tray | Specialised inference processing |
| BlueField 4 storage tray | Storage, networking, and security offload |
NVIDIA describes this approach as extreme co design. Instead of optimising each processor independently, the company designs compute, networking, storage, cooling, and software as one system.
This matters because large AI models often spend considerable time moving data between accelerators. Faster chips provide limited benefit if networking or memory becomes the main bottleneck.
Early DeepSeek R1 testing shows a major efficiency gain
The reported DeepSeek R1 result compares Vera Rubin with the GB200 NVL72 platform.
| Platform | Reported token throughput | Stated power level |
|---|---|---|
| GB200 NVL72 | 80,000 tokens per second | 150MW |
| Vera Rubin NVL72 | 800,000 tokens per second | 150MW |
| Claimed improvement | 10 times | Same power level |
The comparison suggests Vera Rubin can complete substantially more inference work without increasing the overall power budget.
Token throughput is an important metric for companies running AI services because it directly affects how many responses a data centre can generate. Higher throughput can improve revenue potential while reducing the energy cost per token.

However, token throughput alone does not describe the full experience. Time to first token, response latency, model quality, and consistency also matter. A system that produces many tokens but responds slowly may not be suitable for interactive applications.
NVLink 6 improves communication inside large AI systems
NVIDIA says NVLink 6 provides twice the performance of its previous scale up technology while reducing latency.
The company also claims a tenfold increase in packet rate, three times lower latency, and 130 TFLOPS of in network compute. These improvements help accelerators exchange information more efficiently during large model training and inference.
The networking stack extends beyond the rack. Spectrum X Ethernet Photonics is designed to improve scale out communication, while Spectrum XGS Ethernet supports connections across multiple data centre sites.
| Networking improvement | Reported gain |
|---|---|
| NVLink 6 performance | 2 times |
| Packet rate | 10 times higher |
| Latency | 3 times lower |
| In network compute | 130 TFLOPS |
| Scale out RDMA bandwidth | 1.6 times higher |
| Multi site performance | 1.9 times higher |
Efficient networking becomes increasingly important as AI deployments expand from dozens of accelerators to thousands. Communication delays can quickly reduce scaling efficiency when each processor must wait for data from another part of the system.
The rack design removes more cables and fans
NVIDIA has also redesigned the physical construction of the Vera Rubin racks.
The company says it has reduced the use of cables, housings, and individual fans in favour of a more modular liquid cooled system. This should simplify installation, improve serviceability, and help manage the heat generated by dense compute hardware.
Liquid cooling is becoming more common in high performance AI data centres because traditional air cooling struggles with the power density of modern accelerator racks.
A cleaner rack design can also improve reliability by reducing the number of physical connections and moving parts. However, liquid cooling requires specialised facility infrastructure and careful maintenance.
First systems are beginning to power on
The first Vera Rubin systems are now being installed and validated by major cloud infrastructure providers.
Early deployment does not mean the platform is already available at full scale. Customers must test reliability, software compatibility, networking performance, cooling, and model behaviour before using the systems broadly.
Performance is also likely to improve after deployment. NVIDIA has already shown that software updates can deliver large gains on existing Blackwell hardware without changing the physical system.
Vera Rubin will therefore launch with strong early figures, but its long term performance may depend just as much on software optimisation as on the new chips.
Vera Rubin strengthens NVIDIA’s rack scale strategy
The reported tenfold increase over GB200 demonstrates why NVIDIA is moving beyond selling individual accelerators.
The company now competes through complete AI factories that include processors, memory, networking, storage, cooling, and a mature software stack. This integrated approach can reduce deployment complexity for customers, but it also makes them more dependent on NVIDIA’s ecosystem.
AMD and other competitors are pursuing similar rack scale designs, including systems built around high capacity HBM memory and open Ethernet networking.
Independent benchmarks will be necessary to determine how Vera Rubin performs across different models and frameworks. The early DeepSeek R1 result is impressive, but it represents one workload under a specific configuration.
Still, reaching 800,000 tokens per second at the same stated power level as a GB200 system shows the scale of improvement NVIDIA is targeting. Vera Rubin is designed to increase AI output without requiring a proportional rise in electricity use, which may become its most important advantage as data centre power limits continue to tighten.



Discussion (0)
Be the first to comment.