A new memory technology called HBF, or High-Bandwidth Flash, is starting to attract attention in the AI hardware world. It promises much higher capacity than today’s HBM memory, but NVIDIA reportedly is not planning to use it yet. Instead, Google is said to be one of the early major customers as sampling begins later this year.
HBF is being developed as a way to sit between high-speed HBM and traditional NAND flash storage. It uses stacked NAND layers connected with TSVs, which is similar in idea to how HBM stacks memory. The key difference is capacity. While HBM stacks today usually offer tens of gigabytes, HBF is expected to scale up to as much as 4TB per stack.
That capacity could be very useful for AI, especially inference workloads. Modern AI systems often need to hold large amounts of data close to the processor, and memory limits can become a major bottleneck. HBF may not match HBM in raw speed, but its large capacity could help with workloads where storing more data near the compute chip matters.
HBF could help AI companies handle larger memory needs, but NVIDIA seems happy with HBM and faster SSDs for now
According to the report, NVIDIA is not interested in adopting HBF in the near term. The company is instead expected to keep relying on HBM for its AI accelerators. NVIDIA also reportedly believes that enterprise SSDs can help solve some capacity and speed problems, especially as faster PCIe Gen7 SSDs are developed.
That approach makes sense for NVIDIA. HBM is already deeply tied to its current AI GPU roadmap. It is fast, proven, and supported by the wider packaging ecosystem. Moving to a new memory type too early could add risk, even if HBF looks promising on paper.
Google may see the situation differently. The report says Google is expected to be a major HBF customer as it continues expanding its TPU ecosystem. Google designs its own AI chips, so it may have more room to experiment with memory layouts that fit its internal workloads.
| Memory type | Main strength | Main limitation |
|---|---|---|
| HBM | Very high bandwidth for AI accelerators | Lower capacity per stack than HBF |
| HBF | Much higher capacity, possibly up to 4TB per stack | Slower than HBM and still early |
| Enterprise SSDs | Large storage capacity and improving speed | Farther from the processor than stacked memory |
HBF could also be useful beyond replacing or supporting HBM. The report notes that it may become an option for reducing the need for large amounts of standard DDR or LPDDR memory in some servers. Because it is stacked, it could save board space while still offering high capacity and reasonable bandwidth.
This matters because AI hardware is changing quickly. GPUs were once the main focus, but memory, storage, and CPU bottlenecks are becoming just as important. As AI models and agentic systems grow, companies need more than raw compute power. They need enough memory close to the processor to keep those systems running efficiently.
For now, HBF is still an early technology. SK Hynix is reportedly leading development, with first samples expected in the second half of this year. NVIDIA may not be ready to use it, but Google’s reported interest suggests HBF could still become important in future AI systems.
The bigger story is that the AI memory race is widening. HBM will likely remain the premium choice for top-end accelerators, but HBF could become a serious option where capacity matters more than absolute bandwidth.



Discussion (0)
Be the first to comment.