NVIDIA says continued software optimisation has delivered major performance gains across its Blackwell AI platforms, with the GB300 setting new records in large model training while the older GB200 becomes significantly more efficient.
The company reports that the GB200 NVL72 improved token throughput per megawatt by four times in only a few months while running the DeepSeek R1 0528 model. The hardware configuration remained unchanged, which means the increase came primarily from software, framework, and system level improvements.
NVIDIA carried out more than 250,000 simulated configurations and spent approximately 1.4 million GPU hours testing possible optimisations. The company says this work produced 38 major improvements, with more than 90 percent of them applicable to other AI models.
These figures show that AI accelerator performance does not depend only on new chips. Software updates, communication libraries, memory management, scheduling, and framework support can continue raising performance after hardware has already reached data centres.
GB300 delivers record performance for DeepSeek V3 training
The newer GB300 Blackwell Ultra platform achieved 1,648 TFLOPS per GPU while pretraining the 671 billion parameter DeepSeek V3 model across 256 GPUs.
That result is nearly three times higher than the 606 TFLOPS per GPU reported for the GB200 NVL72 in the same comparison. NVIDIA says this allows the GB300 system to complete the same training workload with considerably less hardware.
| Blackwell result | Reported performance |
|---|---|
| GB300 Megatron Core training | 1,648 TFLOPS per GPU |
| GB200 comparison result | 606 TFLOPS per GPU |
| GB300 improvement over GB200 | Nearly 3 times |
| Earlier GB300 result | 1,088 TFLOPS per GPU |
| Latest GB300 result | 1,648 TFLOPS per GPU |
| Improvement over six months | Around 1.5 times |
The GB300 result also represents a substantial increase over its own earlier performance. Megatron Core improved from 1,088 TFLOPS per GPU in November 2025 to 1,648 TFLOPS per GPU by June 2026.
This means the platform gained around 50 percent more delivered throughput without requiring a completely new hardware generation.
GB200 efficiency increased fourfold through software work
NVIDIA also highlighted the progress made with the GB200 NVL72.
While running DeepSeek R1 0528 with a 1,000 token input and 1,000 token output workload, the same GB200 configuration delivered four times more token throughput per megawatt than it had only a few months earlier.

Performance per watt is particularly important for large AI data centres. Power availability has become one of the main limits on AI expansion, with some deployments consuming tens or hundreds of megawatts.
A fourfold increase in throughput per megawatt allows customers to process more requests without installing additional power infrastructure. It can also reduce the cost of running inference workloads at scale.
The improvement does not mean the GPU itself became four times faster in every task. The result applies to a specific workload and configuration, but it demonstrates how much unused performance can remain available until software is fully tuned.
PyTorch and JAX also receive major gains
NVIDIA is working closely with the PyTorch and JAX communities to improve performance across several training frameworks.
Using TorchTitan, PyTorch’s native training stack, the GB300 NVL72 delivered a sixfold improvement over an unoptimised baseline while training DeepSeek V3.
Performance increased from 199 TFLOPS per GPU to 1,197 TFLOPS per GPU after optimisation.
JAX showed an even larger improvement. The platform reached 1,025 TFLOPS per GPU and up to 4,082 tokens per second per GPU. That is almost ten times the 418 tokens per second per GPU recorded in January 2026.
| Framework | Earlier or baseline result | Optimised result | Improvement |
|---|---|---|---|
| Megatron Core | 1,088 TFLOPS per GPU | 1,648 TFLOPS per GPU | 1.5 times |
| TorchTitan | 199 TFLOPS per GPU | 1,197 TFLOPS per GPU | 6 times |
| JAX | 418 tokens per second per GPU | 4,082 tokens per second per GPU | Nearly 10 times |
These gains are important because AI developers do not all use the same software stack. Strong performance across Megatron Core, PyTorch, and JAX gives customers more flexibility when selecting tools for training large models.
Scaling remains efficient across 1,024 GPUs
Large model training requires thousands of accelerators to work together efficiently. Performance can fall as more GPUs are added because communication overhead increases.
NVIDIA says the GB300 NVL72 maintains strong scaling from 256 to 1,024 GPUs. Megatron Core reportedly reaches 98.5 percent scaling efficiency at 1,024 GPUs, while TorchTitan and JAX remain close to 97 percent.
The 800 Gb per second scale out networking used within and between NVL72 racks plays an important role in maintaining this efficiency.
Strong scaling means customers can add more GPUs without losing a large amount of performance to communication delays. This is essential for training massive mixture of experts models, where data must move quickly between different parts of the system.
Blackwell will remain important even as Rubin arrives
NVIDIA is beginning to introduce its next generation Vera Rubin platform, but Blackwell systems will remain widely used for years.
Data centres have already invested heavily in GB200 and GB300 hardware. Continued software improvements can extend the useful life of those deployments and improve their financial return.
The latest results also show why hardware comparisons based only on launch day benchmarks can be misleading. AI platforms often improve significantly after deployment as compilers, frameworks, networking libraries, and model specific optimisations mature.
GB300 currently leads the reported Blackwell training results, while GB200 continues to gain efficiency through software. Together, they show that NVIDIA is treating optimisation as an ongoing process rather than shifting all attention immediately to the next architecture.



Discussion (0)
Be the first to comment.