NVIDIA is expanding local AI support across its RTX and DGX platforms with new performance optimizations and simpler setup tools for systems equipped with at least 24GB of GPU memory.
The updates cover both inference speed and usability. NVIDIA says recent work with the llama.cpp and vLLM communities can improve local AI performance by up to 1.9x in some workloads, while new one click setup options are coming to tools such as Hermes Agent, OpenClaw, and Perplexity Portable Computer.
The changes are aimed at making local AI easier to run without requiring manual model downloads, configuration, or extensive tuning.
| Platform or tool | Reported improvement or feature |
|---|---|
| GeForce RTX 5090 with llama.cpp | Up to 50% faster on Qwen3.6 27B |
| GeForce RTX 5090 with llama.cpp | Up to 90% faster on Qwen3.6 35B |
| RTX PRO 6000 Blackwell with vLLM | Up to 20% faster |
| DGX Spark with vLLM | Up to 1.4x faster on Qwen3.6 27B |
| DGX Spark | Up to 20% faster on DeepSeek v4 Flash |
| Minimum VRAM for new simplified local AI support | 24GB |
| Supported operating systems | Windows and Linux, depending on app |
RTX 5090 Sees the Biggest Reported Gains
NVIDIA says its latest llama.cpp optimizations can substantially improve token throughput on GeForce RTX systems.
On the RTX 5090, Qwen3.6 27B reportedly gains up to 50% more token throughput, while Qwen3.6 35B can improve by as much as 90%.
The company says these results were measured using SpeedBench Coding 8K Throughput with AIPerf.
The gains come from continued software and kernel optimization rather than new hardware.
NVIDIA points to faster prefill, improved speculative decoding, and more efficient inference paths as the main contributors.
vLLM Improvements Extend to RTX PRO and DGX Spark
The RTX PRO 6000 Blackwell also benefits from the latest vLLM work.
NVIDIA reports up to 20% higher performance across the tested Qwen3.6 models.
DGX Spark systems see larger gains in some cases, including up to 1.4x higher performance on Qwen3.6 27B and around 20% faster inference on DeepSeek v4 Flash.

The vLLM improvements include new XQA attention kernels in FlashInfer and other backend changes intended to make local inference more efficient.
Actual performance will still depend on model size, quantization, context length, prompt structure, and the exact GPU configuration.
One Click Local AI Is Coming to 24GB NVIDIA GPUs
The second part of NVIDIA's announcement focuses on simplifying local AI setup.
Systems with at least 24GB of NVIDIA GPU memory will gain streamlined support for several agent platforms.
Perplexity Portable Computer is expected to become available on compatible RTX GPUs running Windows or Linux in September.
The app is designed to package models, orchestration tools, and local workflows into a simpler setup.
It can run tasks locally without consuming cloud credits, while still allowing selected parts of a workflow to be sent to cloud based frontier models when needed.
Perplexity says the app asks for permission before sending content to the cloud.
Hermes Agent Will Detect the GPU Automatically
Hermes Agent is also getting a simpler local setup process.
The upcoming configuration system is expected to automatically detect the installed NVIDIA GPU, choose a suitable model and configuration, and launch it through integrated llama.cpp support.
That removes several steps that normally require manual model selection, downloading, and tuning.
Once running locally, Hermes can continue using tools, maintaining task context, remembering information across sessions, and creating reusable skills.
NVIDIA says the local setup will work across RTX PCs, RTX PRO workstations, and DGX Spark systems.
OpenClaw Setup on Windows Is Also Being Simplified
NVIDIA, Microsoft, and OpenClaw are also working together on a Windows application that reduces the setup effort for local models.
The OpenClaw Windows App is designed for RTX systems with at least 24GB of VRAM.
It can help configure the Windows Subsystem for Linux environment and prepare a compatible local AI model automatically.
That could make OpenClaw more accessible to people who want local agents but do not want to manually configure Linux environments, dependencies, and inference software.
Local AI Is Becoming More Software Driven
The latest changes show that performance improvements for local AI are increasingly coming from software as much as hardware.
Kernel tuning, attention optimizations, better speculative decoding, and easier setup can all improve the experience without replacing the GPU.
For people with high memory RTX cards, the new tools may reduce the gap between installing a local model and actually using it for practical workflows.
NVIDIA has not given one release date for every feature, but the Perplexity Portable Computer rollout is expected in September, while simplified Hermes and OpenClaw setup is listed as coming soon.



Discussion (0)
Be the first to comment.