Qualcomm has detailed a major upgrade to its Hexagon NPU that is designed to run much larger AI models directly on future smartphones.
The redesigned NPU will support Mixture of Experts models with up to 30 billion total parameters while activating only about 3 billion parameters for each token generation step. Qualcomm says this approach can reduce memory bandwidth and power demands while keeping large amounts of model knowledge available on the device.
The new Hexagon architecture will join Qualcomm's next generation Oryon CPU and Adreno Neural Fusion graphics technology as part of its upcoming flagship Snapdragon mobile platform.
Hexagon NPU gets more shared memory
One of the largest hardware changes is a 50% increase in shared memory compared with the NPU used in Snapdragon 8 Elite Gen 5.
More shared memory allows the accelerator to keep additional model state, context, and KV cache data close to its compute units instead of repeatedly moving that information through slower parts of the memory system.
| Feature | Next generation Hexagon NPU |
|---|---|
| Shared memory increase | 50% versus Snapdragon 8 Elite Gen 5 |
| Maximum reported model size | Up to 30 billion parameters with MoE |
| Active parameters per token | About 3 billion |
| Maximum context | Up to 32,000 tokens |
| New accelerator | Element Accelerator |
| Main focus | Generative and agentic AI |
| Target devices | Future flagship smartphones |
| Full platform reveal | Snapdragon Summit, September 22 |
Qualcomm also says dedicated KV cache acceleration will support context windows of up to 32,000 tokens.
That larger context capacity could help local AI systems work with longer conversations, documents, and multi step tasks without relying as heavily on cloud processing.
Element Accelerator targets transformer workloads
The redesigned NPU introduces a new component called the Element Accelerator.
It combines vector extensions for high throughput mathematical operations with scalar extensions for execution logic, decision making, and tool routing.
Those functions are particularly relevant to transformer based generative AI and agentic systems, where models may need to process information, choose between tools, and manage several operations within a single task.
Qualcomm says the accelerator is intended to improve both response speed and efficiency.
The company is positioning the architecture around persistent multimodal AI workloads that can remain active on a smartphone without placing excessive pressure on battery life.
Mixture of Experts makes larger models more practical
Running a conventional 30 billion parameter dense model locally on a phone would create major memory and bandwidth challenges.
Qualcomm is therefore focusing on Mixture of Experts, or MoE, architectures.

Instead of activating every parameter for every request, an MoE model contains several specialized groups of parameters and dynamically selects only the most relevant ones for a particular input.
In Qualcomm's example, a model can contain 30 billion total parameters while using roughly 3 billion active parameters during each token generation step.
This reduces the amount of computation required at any moment.
It can also lower memory bandwidth requirements, which is particularly important on mobile devices where power consumption and thermal limits are much tighter than on desktop PCs or servers.
Flash storage will help manage model experts
Qualcomm also plans to use intelligent flash to memory management and caching.
This allows parts of a large model to remain in storage until they are required, rather than loading every expert into active memory at the same time.
The system can then move frequently needed experts into faster memory while managing less active model components from flash storage.
That technique is intended to make larger local models possible without requiring smartphone memory capacities that would otherwise be impractical.
Actual performance will depend heavily on model design, memory configuration, storage speed, and software optimization.
Qualcomm is building around local agentic AI
The NPU announcement follows earlier disclosures about Qualcomm's upcoming mobile architecture.
The company has already discussed a next generation Oryon CPU capable of reaching 5GHz and an Adreno Neural Fusion graphics architecture designed to combine graphics and AI processing.
Together, these components form Qualcomm's broader strategy for running more advanced AI locally rather than sending every request to a cloud service.
Local processing can potentially improve latency and privacy because some information can remain on the device, although actual behavior will depend on individual applications and services.
Qualcomm is expected to provide more information about the complete platform at Snapdragon Summit, which begins on September 22, 2026.



Discussion (0)
Be the first to comment.