NVIDIA has expanded local AI support across its RTX and DGX systems with new model optimizations, faster video generation, broader open model compatibility and new tools for clustering DGX Spark systems.
The update covers several areas at once. Meta’s Muse Glimmer can reportedly exceed 200 tokens per second on an RTX 5090, while LTX 2.5 video generation gains up to twice the performance and 40 percent lower memory use on supported NVIDIA hardware.
NVIDIA is also adding support for a wider group of open models, including DeepSeek V4 Flash, MiniMax H3, Cosmos 3 Edge and Wan Animate 2.
NVIDIA AI platform update highlights
| Feature | Detail |
|---|---|
| Muse Glimmer | More than 200 tokens per second on RTX 5090 |
| LTX 2.5 | Up to 2x performance improvement |
| LTX 2.5 memory use | Up to 40% lower |
| DGX Spark | New Sync Cluster Assistant |
| Resource monitoring | New Sync Resource Monitor |
| Browser support | Native ARM64 Linux build of Google Chrome |
| Open model support | Expanded across RTX and DGX |
| Unsloth Desktop | Local inference, fine tuning, diffusion, agents and code execution |
Muse Glimmer targets local agent workloads
Meta’s Muse Glimmer is a 30 billion parameter open weight model designed for coding and agent based workloads.
On a single RTX 5090, NVIDIA says the model can run at more than 200 tokens per second.
The model is designed for tasks that need sustained local processing, including working with private files, executing multistep workflows and interacting with tools.
Its local execution model also allows sensitive information such as documents, messages, API credentials and authentication tokens to remain on the device rather than being sent to a remote server.
Muse Glimmer also supports long running tasks that can be broken into multiple stages and resumed later with context preserved.
LTX 2.5 improves local video generation
LTX 2.5 is another major part of the update.
The model adds multishot generation and generative editing while introducing a new diffusion video decoder alongside the existing variational autoencoder decoder.

The second decoding path is intended to improve final video quality.
On NVIDIA RTX GPUs, DGX Spark and DGX Station, LTX 2.5 can deliver up to twice the performance while reducing memory use by as much as 40 percent.
The RTX PRO 6000 Blackwell is cited with a 20 percent performance improvement and 40 percent lower memory use during video generation.
NVIDIA is also supporting related tools including NVFP4, FastVideo and ComfyUI for text, image and video based generation workflows.
More open models are now accelerated
NVIDIA is broadening support for several open models across its client AI platforms.
The newly supported list includes:
| Model | Parameter count |
|---|---|
| Cosmos 3 Edge | 4 billion |
| MiniMax H3 | 33 billion |
| Laguna S 2.1 | 118 billion |
| DeepSeek V4 Flash | 284 billion |
| Wan Animate 2 | 14 billion |
| Inkling Small | 276 billion |
The update also includes Unsloth Desktop, an open source application that supports local inference, fine tuning, image and video diffusion, agent integrations, web research and code execution.
NVIDIA says RTX hardware can deliver large performance advantages in some of these workloads.
For Wan Animate 2, the RTX 5090 is claimed to provide up to 26 times the performance of an Apple M3 Ultra, while the RTX PRO 5000 Blackwell is cited at up to 16 times faster.
DGX Spark gains easier clustering
DGX Spark is receiving its own group of updates.
The new NVIDIA Sync Cluster Assistant is designed to make it easier to connect two or more DGX Spark systems into a single cluster.
The software handles networking, workload routing and system health monitoring, allowing multiple systems to combine memory and compute resources for larger workloads.
NVIDIA is also adding Sync Resource Monitor, which provides current and historical CPU and GPU usage across either one machine or an entire DGX Spark cluster.
Google Chrome is also coming to DGX Spark as a native ARM64 Linux build and can be installed directly through the DGX Dashboard.
Local AI remains a major focus for RTX and DGX
The latest updates show NVIDIA continuing to expand local AI beyond simple chatbot inference.
RTX and DGX systems are being positioned for coding, local agents, video generation, fine tuning and large open models, while DGX Spark is gaining tools that make multiple systems easier to combine.
The strongest headline figures are the more than 200 tokens per second claimed for Muse Glimmer on an RTX 5090 and the performance and memory improvements in LTX 2.5.
Together, the changes give developers and creators more options for running demanding AI workloads locally without depending entirely on cloud infrastructure.



Discussion (0)
Be the first to comment.