Arm AGI CPU Packs 136 Neoverse V3 Cores and More Than 100 Billion Transistors for Agentic AI

news
Arm AGI CPU Packs 136 Neoverse V3 Cores and More Than 100 Billion Transistors for Agentic AI

Arm has provided a deeper look at its AGI CPU, a high performance processor designed around agentic AI workloads and built with up to 136 Neoverse V3 cores.

The chip uses a dual chiplet design manufactured on TSMC’s 3nm process, with each chiplet containing more than 50 billion transistors. Arm says the processor is intended to handle the CPU intensive parts of AI systems, including orchestration, tool execution, planning, reasoning and data processing.

Demand expectations are also rising. Arm has increased its forecast for AGI CPU customer demand to $2 billion across 2027 and 2028, up from an earlier $1 billion estimate.

Arm AGI CPU specifications

FeatureDetail
ArchitectureArm Neoverse V3
Maximum core count136 cores
ProcessTSMC 3nm
DesignDual chiplet
TransistorsMore than 50 billion per chiplet
Maximum frequencyUp to 3.7 GHz
L2 cache2MB per core
Maximum TDP300W
MemoryDDR5 8800
Memory capacityUp to 6TB per chip
PCIe96 PCIe Gen6 lanes
ExpansionCXL 3.0
Target workloadsAgentic AI and cloud infrastructure

Arm says each AGI processor can reach frequencies of up to 3.7 GHz while operating within a maximum 300W TDP. Each Neoverse V3 core receives 2MB of private L2 cache.

Agentic AI puts more pressure on CPUs

Arm argues that CPUs are becoming increasingly important as AI moves beyond simple prompt and response workloads.

AI agents often need to coordinate tools, execute code, move data and perform reasoning between model calls. Accelerators may handle LLM prefill and decode stages, but CPUs remain responsible for much of the surrounding workflow.

Arm says token consumption from AI agents could increase 24 times by 2030. It also estimates that enterprise automation workloads can split tasks roughly 50 percent between CPUs and accelerators, while software development can lean even more heavily on CPUs.

Different workloads stress different parts of the processor. Tool use depends on IPC, I/O and concurrency, while orchestration emphasizes utilization and parallel activity. Reasoning places more pressure on branch prediction, cache performance, latency and memory bandwidth.

Neoverse V3 uses a wide out of order design

The Neoverse V3 core is based on Armv9.2 and uses a 10 wide front end.

Its architecture includes a large out of order execution window, third generation prefetching, advanced branch prediction and a low latency private L2 cache.

The core includes eight integer ALUs, three branch units, dual 128 bit SIMD pipelines and 64KB instruction and data caches. Its 2MB L2 cache offers a 10 cycle load to use latency and supports ECC and parity protection.

Arm is clearly prioritizing per core efficiency and latency rather than simply increasing the total number of cores.

Two chiplets contain up to 140 physical cores

The AGI CPU physically contains up to 70 cores on each of its two chiplets, for 140 cores in total.

Arm disables four cores for yield reasons, leaving a maximum usable configuration of 136 cores.

Each chiplet includes 64KB instruction and data caches per core, 2MB of private L2 per core and a shared system level cache allocation of 1MB per core. The chiplets are tied together through a coherent interconnect and UCIe based die to die subsystem.

Memory controllers sit around the edges of the chiplets, while the CPU core complexes occupy the center.

Memory and I/O are designed for low latency

Arm integrates memory and I/O functionality closely with the compute dies.

The company claims sub 100 nanosecond memory latency and supports DDR5 8800 memory, with up to 6TB of capacity per processor.

The CPU also provides 96 PCIe Gen6 lanes, CXL 3.0 and AMBA CHI extension links.

Arm says each core can receive around 6GB/s of memory bandwidth, which is important for AI orchestration workloads where cache misses and memory latency can quickly limit performance.

Rack scale systems can reach 8,160 cores

Arm is planning both reference servers and partner systems around the AGI CPU.

Its own designs include 1U dual node and 2U dual processor servers, while companies including Supermicro, ASRock and Lenovo are preparing their own implementations.

Some rack scale configurations can reach 8,160 CPU cores across 30 1U servers in a 36kW power envelope.

Arm claims those systems can deliver more than twice the performance per rack of an x86 based system operating within the same power budget.

That performance claim comes from Arm and will need independent testing once production systems become available.

Arm expects AI infrastructure demand to keep growing

Arm says it has now shipped 1.5 billion Neoverse cores into data centers, including 500 million during the past nine months.

The company also claims more than 60 percent share as the host CPU architecture in AI server platforms and says over 120,000 companies have deployed Arm based systems in the cloud.

The AGI CPU represents Arm’s effort to turn that architectural presence into its own dedicated AI focused processor platform.

With 136 Neoverse V3 cores, a 300W power target, large memory capacity and PCIe Gen6 connectivity, the chip is positioned directly for the growing CPU requirements surrounding agentic AI rather than traditional accelerator workloads alone.

Discover: News

Discussion (0)

Be the first to comment.