Modders have managed to run a hybrid FP8 and NVFP4 version of NVIDIA’s DLSS 5 Neural Rendering model on GeForce RTX 50 series graphics cards, but the early results show only a very small performance improvement.
The experiment is designed to reduce the processing cost of DLSS 5 by moving parts of the neural workload from FP8 to NVFP4. Blackwell GPUs include fifth generation Tensor Cores with native support for NVFP4, making the format an obvious candidate for reducing inference cost.
The practical gains, however, are limited so far.
Testing showed that the complete DLSS 5 Neural Rendering pass at 4K became only around 1% to 2% faster compared with the original FP8 implementation. That figure applies specifically to the neural rendering workload and should not be interpreted as a 1% to 2% increase in total game performance.
| Test detail | Result |
|---|---|
| GPU generation | GeForce RTX 50 series |
| Neural formats | Hybrid FP8 and NVFP4 |
| Neural pass improvement | Around 1% to 2% |
| Baldur’s Gate 3 hybrid result | 55.16 rendered FPS |
| Baldur’s Gate 3 FP8 result | 54.57 rendered FPS |
| Main limitation | Small and potentially inconsistent gains |
NVFP4 does not automatically make DLSS 5 much faster
NVFP4 uses lower numerical precision than FP8 and can reduce memory requirements for neural workloads.
On paper, that can provide better inference performance on hardware designed to accelerate the format. The challenge is that a neural network usually needs to be properly quantized and optimized for lower precision before large gains can be achieved.
Simply converting parts of an existing FP8 model to NVFP4 does not guarantee a major improvement.
The current implementation also continues to use FP8 for model shapes that are not supported by the NVFP4 path. This mixed approach explains why the project is described as a hybrid implementation rather than a complete conversion.
A later version of the project made the hybrid path the recommended option for Blackwell cards after additional optimization.
One 120 second Baldur’s Gate 3 test recorded 55.16 rendered frames per second with the hybrid model, compared with 54.57 FPS using FP8.
That difference is small enough that it should not yet be treated as a repeatable performance advantage. More testing across different games, scenes and GPUs would be required before drawing a firm conclusion.
DLSS 5 remains expensive to run
The experiment highlights a broader issue with DLSS 5.
NVIDIA’s latest Neural Rendering model offers more advanced image reconstruction, but it also places a heavier computational load on the GPU than previous DLSS implementations. Reducing that cost is therefore an important area for both official development and community experimentation.

NVIDIA has already improved DLSS 5 substantially since its earlier demonstrations. The company has said the model is now around five times faster than its original GTC 2026 version, and further optimization is planned.
Official RTX 40 series support is also expected, which will make efficiency even more important because older GPUs have less AI processing capability than Blackwell based RTX 50 hardware.
For now, the NVFP4 experiment shows that lower precision alone is not enough to produce a major breakthrough. The 1% to 2% reduction in neural pass time is technically useful, but it remains modest.
Larger gains will likely require deeper model optimization, better quantization and changes designed specifically around the strengths of NVFP4 rather than converting parts of an FP8 model after the fact.



Discussion (0)
Be the first to comment.