d Matrix Raptor 3D DRAM Targets 100 TB/s Bandwidth at 0.37 pJ per Bit for AI Inference

news
d Matrix Raptor 3D DRAM Targets 100 TB/s Bandwidth at 0.37 pJ per Bit for AI Inference

d Matrix has detailed its Raptor 3D DRAM architecture, a stacked memory design aimed at AI inference workloads where bandwidth and power consumption are becoming major limits.

The company positions Raptor between SRAM and HBM. SRAM can deliver extremely high bandwidth and low energy per bit, but its capacity and cost make it difficult to scale. HBM offers much larger capacity, but increasing bandwidth can require substantial power and complex packaging.

Raptor attempts to combine high bandwidth with more practical capacity by placing compute logic directly on top of custom DRAM and using very short vertical interconnects.

Raptor 3D DRAM specifications

FeatureDetail
Memory architecture3D stacked DRAM
Logic processTSMC 4nm
Stacking methodFace to face
Interconnect pitch36 microns
Capacity32GB per 1 high stack
Bandwidth100 TB/s
Measured energy0.37 pJ per bit
Tensor engines256 per chiplet
Logic power densityAbout 0.5 W/mm²
Refresh interval4 ms
Error correctionReed Solomon T=2 plus CRC

d Matrix says a single Raptor 3D DRAM configuration can deliver around 100 TB/s of bandwidth with measured energy consumption of 0.37 pJ per bit.

The design uses a TSMC N4 logic die stacked face to face above custom DRAM using a 36 micron connection pitch.

Removing the PHY reduces data movement overhead

A major part of the efficiency gain comes from eliminating some of the long distance interfaces associated with conventional HBM.

Raptor places logic directly above the DRAM, allowing the compute engines to communicate through short vertical connections rather than moving data across a conventional high speed PHY.

d Matrix claims that a 3D DRAM design with four or fewer memory layers can provide much higher bandwidth while consuming around one tenth of the energy of HBM in comparable data movement scenarios.

The company also pitch matches the memory banks to 256 tensor engines per chiplet. This is intended to minimize unnecessary data movement between memory and compute.

100 TB/s still requires significant I/O power

Raptor does not eliminate the power problem entirely.

At 0.37 pJ per bit, moving data at 100 TB/s still requires roughly 300W of I/O power.

That remains substantial, but d Matrix argues it is much lower than what an HBM based design would require to reach similar bandwidth.

The company estimates that a hypothetical 100 TB/s HBM solution operating at around 2.4 pJ per bit would consume approximately 1.92kW for the HBM memory alone.

Thermal management requires faster refresh

Stacking logic directly on DRAM creates thermal challenges because DRAM becomes less reliable as temperature rises.

Raptor places the logic die on top so cooling hardware can contact the hotter compute layer directly.

The DRAM is designed to operate at junction temperatures up to 105°C, but it compensates by refreshing eight times more frequently than conventional DRAM, approximately once every 4 milliseconds.

d Matrix says the resulting bandwidth loss from refresh is only 1.37%.

Reliability is handled with small banks and strong ECC

The design uses small microbanks containing 1,366 rows and about 5.33MB each.

To improve manufacturing yield, roughly 8% to 9% spare banks are included. Raptor also uses Reed Solomon T=2 error correction on the logic die, allowing it to correct two symbol errors across a 128 byte block, alongside CRC protection.

These techniques are intended to make dense 3D memory more practical despite the added complexity of stacking logic and DRAM.

d Matrix claims large gains over HBM4

d Matrix compares Raptor against an NVIDIA Rubin R200 platform using HBM4.

With effective bandwidth utilization of around 83% to 85%, the company claims Raptor provides 23.4 times higher bandwidth per square millimeter and 13.5 times better power efficiency measured in milliwatts per GB/s.

Those figures are vendor supplied comparisons and will need independent validation.

The broader goal is clear: Raptor is designed for AI inference workloads, particularly decode stages where models repeatedly access large KV caches and can become limited by memory bandwidth rather than raw compute.

If the architecture scales as d Matrix expects, 3D DRAM could offer a middle ground between the capacity of HBM and the extreme bandwidth of SRAM for future AI accelerators.

Discover: News

Discussion (0)

Be the first to comment.