AMD is partnering with Cerebras to combine its Helios rack scale AI platform with the Cerebras Wafer Scale Engine, creating a system designed for high throughput and low latency inference.
The companies say the integrated platform can deliver up to five times more tokens per second per watt than Cerebras hardware operating alone. The goal is to divide AI workloads between two specialised systems rather than forcing one type of accelerator to handle every stage.
Helios will process prompts, large context windows and high volume workloads using AMD Instinct accelerators. Cerebras hardware will focus on the memory intensive decode stage, where models generate tokens and fast response times become especially important.
The combined system is expected to become available during the second half of 2026. AMD is positioning the collaboration as a flexible alternative to NVIDIA’s growing investment in Groq and its specialised Language Processing Unit technology.
| Component | Main role |
|---|---|
| AMD Helios | Prompt processing, large contexts and scalable throughput |
| Cerebras Wafer Scale Engine | Low latency decoding and token generation |
| AMD Instinct MI455X | High capacity AI training and inference |
| Cerebras on chip SRAM | Fast data access without conventional external memory bottlenecks |
| Claimed combined gain | Up to five times more tokens per second per watt |
| Expected availability | Second half of 2026 |
Helios and Cerebras will handle different parts of inference
Modern AI inference contains several stages with different hardware requirements.
The first stage processes the user’s request and any associated context. Long prompts, documents and large context windows require substantial compute power and memory capacity.
The second stage generates the response one token at a time. This decode process is often limited by memory access and latency rather than raw mathematical throughput.
AMD plans to use Helios as the main scalable engine for processing large prompts and high volume requests. The platform combines Instinct MI455X accelerators, Zen 6 EPYC processors, networking and large amounts of HBM4 memory.
Cerebras will handle the parts of inference that benefit from extremely fast on chip memory and low latency communication.
By assigning each stage to the hardware best suited to it, the companies expect to improve both response times and overall efficiency.
This could be especially useful for coding assistants, AI agents and interactive services where customers expect fast responses even when models are large.
Cerebras uses an entire wafer as one processor
The Cerebras Wafer Scale Engine follows a very different design from a conventional GPU.
Instead of cutting a silicon wafer into many individual chips, Cerebras connects a large portion of the wafer into one enormous processor. This allows hundreds of thousands of AI cores and large amounts of SRAM to communicate without relying on external networking between separate accelerators.
The system reportedly includes around 900,000 AI cores, 44GB of on chip SRAM and memory bandwidth reaching approximately 21,000TB per second.
This design can reduce the delays caused by moving data between processors, external memory and networking equipment.
A medium sized model, or a large section of a much bigger model, can remain inside the unified memory structure. That helps reduce communication overhead during token generation.
Cerebras hardware can support both training and inference, giving it more flexibility than some processors built only for one task.
Groq takes a more specialised approach to token generation
NVIDIA’s Groq strategy is based on a different architecture.
Groq’s Language Processing Unit uses specialised Matrix Multiply and Vector processing blocks combined with about 230MB of SRAM on each chip.

The processor does not depend on conventional branch prediction or hardware scheduling. Instead, its compiler plans operations in advance, determining when every calculation and data transfer should occur.
This predictable execution model can produce very fast inference because data moves through the chip according to a fixed schedule.
Large Groq systems distribute a model across hundreds or thousands of individual LPUs. Each chip stores part of the model in its local SRAM.
NVIDIA’s Groq 3 LPX rack reportedly combines 256 Groq 3 accelerators, liquid cooling and up to 315 PFLOPs of inference performance.
The architecture is highly efficient for supported inference workloads, but it is more specialised than a general GPU or the Cerebras Wafer Scale Engine.
AMD is presenting flexibility as a major advantage
The AMD and Cerebras platform is designed to support a broader mix of workloads.
Helios can manage training, prompt processing, long context inference and high throughput deployments. Cerebras adds a specialised engine for low latency token generation.
This allows customers to use the same infrastructure for several stages of an AI workload rather than building a system entirely around one inference accelerator.
Cerebras hardware can also support training, while Groq LPUs are more narrowly focused on inference.
That versatility may appeal to organisations that want to train, fine tune and deploy models using a shared infrastructure strategy.
However, greater flexibility can also increase software and orchestration complexity. The system must move workloads efficiently between Helios and Cerebras without adding enough communication overhead to remove the expected gains.
The claimed five times improvement will therefore depend on how well the software manages both platforms.
The partnership strengthens AMD’s full stack AI strategy
AMD is expanding beyond selling individual CPUs and GPUs.
Helios represents a complete rack scale platform that combines Instinct MI455X accelerators, EPYC processors, networking and software. The Cerebras partnership adds another specialised computing option to that system.
This approach allows AMD to address different parts of the AI market with several hardware types.
The Instinct MI455X provides 432GB of HBM4 memory and up to 40 PFLOPs of FP4 compute. It is designed for large model training and high volume inference.
Cerebras contributes very fast SRAM based processing for workloads where latency and token generation speed matter more than general flexibility.
AMD’s ROCm software environment will remain important because customers need tools that can distribute models and data across the combined infrastructure.
NVIDIA still holds a major advantage through CUDA, its large installed base and tightly integrated AI systems. AMD is attempting to compete through open software, large memory capacity and partnerships with specialised hardware companies.
Efficiency claims will need independent testing
The companies claim up to five times more tokens per second per watt, but the final result will vary by model and deployment.
Performance can depend on model size, prompt length, context window, batch size, precision and how much of the workload runs on each platform.
The comparison also appears to measure the integrated system against Cerebras operating alone rather than directly against NVIDIA’s Groq 3 LPX.
Independent benchmarks will be needed to determine how the AMD and Cerebras platform compares with NVIDIA systems in real production environments.
Power use will also matter. Wafer scale processors, large GPU racks and high speed networking all require significant cooling and electricity.
Even so, the partnership gives AMD a new way to improve inference without designing a separate specialised accelerator internally. Helios can remain the main scalable compute platform, while Cerebras provides the low latency engine for token generation.
The integrated offering is expected during the second half of 2026, with the companies targeting advanced AI services that need both large scale throughput and faster interactive responses.



Discussion (0)
Be the first to comment.