NVIDIA Local AI Gets Up to 1.9x Faster With New llama.cpp and vLLM Optimizations

news
NVIDIA Local AI Gets Up to 1.9x Faster With New llama.cpp and vLLM Optimizations

NVIDIA is expanding local AI support across its RTX and DGX platforms with new performance optimizations and simpler setup tools for systems equipped with at least 24GB of GPU memory.

The updates cover both inference speed and usability. NVIDIA says recent work with the llama.cpp and vLLM communities can improve local AI performance by up to 1.9x in some workloads, while new one click setup options are coming to tools such as Hermes Agent, OpenClaw, and Perplexity Portable Computer.

The changes are aimed at making local AI easier to run without requiring manual model downloads, configuration, or extensive tuning.

Platform or toolReported improvement or feature
GeForce RTX 5090 with llama.cppUp to 50% faster on Qwen3.6 27B
GeForce RTX 5090 with llama.cppUp to 90% faster on Qwen3.6 35B
RTX PRO 6000 Blackwell with vLLMUp to 20% faster
DGX Spark with vLLMUp to 1.4x faster on Qwen3.6 27B
DGX SparkUp to 20% faster on DeepSeek v4 Flash
Minimum VRAM for new simplified local AI support24GB
Supported operating systemsWindows and Linux, depending on app

RTX 5090 Sees the Biggest Reported Gains

NVIDIA says its latest llama.cpp optimizations can substantially improve token throughput on GeForce RTX systems.

On the RTX 5090, Qwen3.6 27B reportedly gains up to 50% more token throughput, while Qwen3.6 35B can improve by as much as 90%.

The company says these results were measured using SpeedBench Coding 8K Throughput with AIPerf.

The gains come from continued software and kernel optimization rather than new hardware.

NVIDIA points to faster prefill, improved speculative decoding, and more efficient inference paths as the main contributors.

vLLM Improvements Extend to RTX PRO and DGX Spark

The RTX PRO 6000 Blackwell also benefits from the latest vLLM work.

NVIDIA reports up to 20% higher performance across the tested Qwen3.6 models.

DGX Spark systems see larger gains in some cases, including up to 1.4x higher performance on Qwen3.6 27B and around 20% faster inference on DeepSeek v4 Flash.

The vLLM improvements include new XQA attention kernels in FlashInfer and other backend changes intended to make local inference more efficient.

Actual performance will still depend on model size, quantization, context length, prompt structure, and the exact GPU configuration.

One Click Local AI Is Coming to 24GB NVIDIA GPUs

The second part of NVIDIA's announcement focuses on simplifying local AI setup.

Systems with at least 24GB of NVIDIA GPU memory will gain streamlined support for several agent platforms.

Perplexity Portable Computer is expected to become available on compatible RTX GPUs running Windows or Linux in September.

The app is designed to package models, orchestration tools, and local workflows into a simpler setup.

It can run tasks locally without consuming cloud credits, while still allowing selected parts of a workflow to be sent to cloud based frontier models when needed.

Perplexity says the app asks for permission before sending content to the cloud.

Hermes Agent Will Detect the GPU Automatically

Hermes Agent is also getting a simpler local setup process.

The upcoming configuration system is expected to automatically detect the installed NVIDIA GPU, choose a suitable model and configuration, and launch it through integrated llama.cpp support.

That removes several steps that normally require manual model selection, downloading, and tuning.

Once running locally, Hermes can continue using tools, maintaining task context, remembering information across sessions, and creating reusable skills.

NVIDIA says the local setup will work across RTX PCs, RTX PRO workstations, and DGX Spark systems.

OpenClaw Setup on Windows Is Also Being Simplified

NVIDIA, Microsoft, and OpenClaw are also working together on a Windows application that reduces the setup effort for local models.

The OpenClaw Windows App is designed for RTX systems with at least 24GB of VRAM.

It can help configure the Windows Subsystem for Linux environment and prepare a compatible local AI model automatically.

That could make OpenClaw more accessible to people who want local agents but do not want to manually configure Linux environments, dependencies, and inference software.

Local AI Is Becoming More Software Driven

The latest changes show that performance improvements for local AI are increasingly coming from software as much as hardware.

Kernel tuning, attention optimizations, better speculative decoding, and easier setup can all improve the experience without replacing the GPU.

For people with high memory RTX cards, the new tools may reduce the gap between installing a local model and actually using it for practical workflows.

NVIDIA has not given one release date for every feature, but the Perplexity Portable Computer rollout is expected in September, while simplified Hermes and OpenClaw setup is listed as coming soon.

Discover: News

Discussion (0)

Be the first to comment.