OpenAI Jalapeno ASIC Reportedly Delivers Higher Inference Efficiency Than NVIDIA Blackwell

news
OpenAI Jalapeno ASIC Reportedly Delivers Higher Inference Efficiency Than NVIDIA Blackwell

OpenAI's first generation Jalapeno AI accelerator is reportedly showing strong inference efficiency, with benchmark results suggesting it can complete 1.5 to 1.9 times more AI work per watt than current NVIDIA Blackwell based systems in selected tests.

The chip combines a reticle sized compute die built on TSMC's N3P process with an N3E I/O chiplet and HBM4 memory. The design reportedly provides 15.4 TB/s of memory bandwidth per package and is paired with OpenAI's own software stack, including its Gluon kernel programming language.

Jalapeno is designed primarily for inference rather than general purpose GPU workloads, so the comparison with NVIDIA hardware should be treated carefully. The reported results come from a specialized benchmark environment and do not mean the ASIC is universally faster across every AI workload.

Jalapeno detailReported specification
Compute processTSMC N3P
I/O chipletTSMC N3E
MemoryHBM4
Memory bandwidth15.4 TB/s per package
Matrix formatMXFP, including MXFP4
Rated TDP700W
Typical active power550W or less
Rack configuration128 Jalapeno ASICs
Rack MXFP4 performanceUp to 1.7 ExaFLOPs
Rack HBM4 capacity27.5 TB

The architecture focuses on reducing data movement

A major part of Jalapeno's design is its use of weight stationary systolic arrays.

Instead of repeatedly moving model weights between memory and compute units, the architecture keeps weights in place while activations flow through the processing array. This can reduce memory traffic and improve efficiency during inference.

The chip also uses MXFP numerical formats, including MXFP4, which compress AI data into lower precision representations. That lowers the amount of memory required and reduces the bandwidth needed to move data through the accelerator.

Alongside the matrix engines, Jalapeno reportedly includes 64 bit scalar cores with out of order execution for control tasks, memory management and scheduling. FP32 and INT32 vector cores handle operations that require higher precision.

Rack scale design reaches 128 accelerators

OpenAI is reportedly building Jalapeno as part of a larger rack scale system rather than as a standalone accelerator.

The host side, called Katsu, contains 16 CPU trays. Each tray uses two AMD EPYC processors, 1.5 TB of DRAM, local NVMe storage and 400G networking.

A corresponding Vindaloo accelerator rack contains 16 trays with eight Jalapeno ASICs each, giving 128 accelerators per rack.

According to the supplied figures, one 128 chip rack can reach up to 1.7 ExaFLOPs of MXFP4 compute and includes 27.5 TB of HBM4 memory.

Benchmark results favor Jalapeno in inference efficiency

Reported InferenceX benchmark results show Jalapeno delivering 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end to end latency than the best available NVIDIA GB200 and GB300 results in the cited tests.

The workloads reportedly included GPT OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T.

Jalapeno is said to provide 13.4 PFLOPs of MXFP4 compute from a single reticle sized die, compared with 17.5 PFLOPs of dense NVFP4 for a single Rubin compute die of similar size on the same process.

Its rated TDP is 700W, though the report says sustained consumption during active AI workloads typically stays at or below 550W.

One important caveat is memory technology. Jalapeno uses HBM4, while the compared GB200 and GB300 systems use HBM3E, giving OpenAI's accelerator an advantage in available bandwidth.

The results also indicate that Jalapeno uses single token prediction rather than speculative decoding or multi token prediction, so the reported performance was not boosted by those techniques.

If these results hold up across broader workloads, Jalapeno could give OpenAI more control over the cost and efficiency of running its models. It could also reduce the company's dependence on general purpose GPU infrastructure for inference, although NVIDIA's broader CUDA ecosystem, software maturity and deployment scale remain major advantages.

Discover: News

Discussion (0)

Be the first to comment.