Alibaba has revealed new details about its next generation Zhenwu V900 AI accelerator, including 216GB of on package memory, native FP8 and FP4 support, and plans to scale clusters to as many as 500,000 chips.
The company also outlined a much larger infrastructure roadmap, targeting more than 20 GW of global data center compute capacity by 2032 while preparing future Qwen AI models with between 4 trillion and 10 trillion parameters.
The Zhenwu V900 is expected to launch in the first quarter of 2027.
Zhenwu V900 will focus on memory and scale
Alibaba says the V900 will include 216GB of on package memory.
That is higher than the 141GB figure cited for NVIDIA's H200 in the supplied material, although memory capacity alone does not determine total accelerator performance.
The exact memory supplier has not been confirmed. There is speculation that the V900 could use HBM3E from CXMT because the reported production schedules appear to overlap, but this remains unverified.
Alibaba is also targeting around three times the performance of its existing Zhenwu M890 accelerator.
| Specification | Zhenwu V900 |
|---|---|
| Expected launch | Q1 2027 |
| On package memory | 216GB |
| Interconnect bandwidth | Around 1.2 TB/s |
| Precision support | FP8 and FP4 |
| Claimed performance gain | Around 3x over M890 |
| Estimated FP16 compute | Around 1.8 PFLOPS |
| Chips per tightly connected accelerator group | Around 1,000 |
| Maximum claimed cluster scale | Up to 500,000 chips |
The 1.8 PFLOPS FP16 estimate comes from applying Alibaba's claimed threefold performance increase to the roughly 0.6 PFLOPS figure associated with the M890.
Alibaba says its ICN Switch fabric will provide around 1.2 TB/s of chip to chip bandwidth.
The company also claims that around 1,000 V900 accelerators can work together as a single logical accelerator through this fabric.
Clusters could reach 500,000 chips
Alibaba's larger ambition goes far beyond individual accelerator groups.
The company says a V900 based cluster could eventually scale to as many as 500,000 chips.
At 216GB of memory per accelerator, that would amount to roughly 108 petabytes of aggregate memory across the cluster.
Reaching that scale would require an enormous supply of memory, packaging capacity, networking hardware and power.
Alibaba has not detailed how quickly such clusters could be deployed or whether the full 500,000 chip configuration represents a near term commercial target.
The company is also developing a broader rack scale stack around its own hardware.
This includes Yitian CPUs, Zhenwu accelerators, ICN networking, Pangu network interface hardware and Zhenyue storage controllers.
That approach suggests Alibaba wants more control over the complete AI infrastructure stack rather than depending entirely on third party accelerator platforms.
Alibaba targets more than 20 GW of compute by 2032
Alibaba also disclosed plans to build more than 20 GW of global data center compute capacity by 2032.
Its current capacity was not disclosed.
An outside estimate from 2025 suggested Alibaba could finish 2026 at around 5 GW, although that figure should be treated as an external estimate rather than a confirmed company number.
If Alibaba were operating at roughly 4 GW to 6 GW by the end of 2026, reaching 20 GW by 2032 would require around 2 GW to 3 GW of additional capacity each year.
Such an expansion would place heavy demands on power generation, data center construction, cooling systems and advanced semiconductor supply.
Qwen 4.5 and Qwen 5.0 could reach 10 trillion parameters
Alibaba also provided an early look at the scale of its future Qwen models.
The upcoming Qwen 4.5 and Qwen 5.0 families are expected to span roughly 4 trillion to 10 trillion parameters.
The company is also exploring recursive self improvement, or RSI.
In this approach, AI systems are used to help identify weaknesses, design experiments, generate data and contribute to the training process of later models.

Alibaba has not provided enough detail to determine exactly how autonomous this process will be or how much human oversight will remain involved.
The scale of the planned models also does not by itself indicate quality, efficiency or real world capability, since architecture, active parameter count, training data and inference design can matter as much as total parameter count.
Qwen Image 2.1 also expands Alibaba's open model push
Alongside its infrastructure announcements, Alibaba has released the weights for Qwen Image 2.1.
The image generation and editing model reportedly uses around 7 billion parameters across 32 single stream DiT layers.
It has also performed strongly in public image model evaluations, particularly among open models.
The broader picture is that Alibaba is building several layers of its AI strategy at once.
It is developing its own accelerators, networking systems, rack scale infrastructure, data center capacity and increasingly large foundation models.
If the Zhenwu V900 arrives on schedule in early 2027, the chip will become an important test of how competitive Alibaba's in house AI hardware can be as the company reduces reliance on external accelerator suppliers.



Discussion (0)
Be the first to comment.