NVIDIA Reportedly Cuts Vera Rubin Memory Capacity as HBM4 Costs Rise

news
NVIDIA Reportedly Cuts Vera Rubin Memory Capacity as HBM4 Costs Rise

NVIDIA is reportedly reducing the memory configuration of its Vera Rubin NVL72 AI rack to control costs as shortages and higher prices continue to affect the memory market.

The adjustment is expected to affect the system’s SOCAMM memory and the memory connected to its Vera CPUs. The HBM4 capacity used by the GPUs is reportedly staying unchanged at 20.7TB per rack.

The standard SOCAMM module capacity could be reduced from 192GB to 96GB. Total Vera CPU memory across the rack may also fall from an earlier estimate of roughly 54TB to 55TB down to about 28TB.

These changes could substantially reduce the cost of each system. Without adjustments, memory was estimated to account for around 29 percent of the VR200 rack’s total bill of materials, above a preferred level of approximately 20 percent.

Vera Rubin NVL72 componentEarlier estimateReported revised configuration
SOCAMM module capacity192GB96GB
Total Vera CPU memoryAround 54TB to 55TBAround 28TB
GPU HBM4 memory per rack20.7TB20.7TB
Estimated memory share of system cost29%Expected to decline
Earlier estimated LPDDR5X costAbout $1.2 millionAs low as $586,000 or $293,000 in reduced configurations

Memory costs could reach almost one third of the rack’s hardware bill

Vera Rubin NVL72 is designed as a rack scale AI platform combining 72 Rubin GPUs with Vera CPUs and large amounts of high bandwidth and system memory.

Memory has become one of the most expensive parts of modern AI infrastructure. Demand from data centres has placed pressure on supplies of HBM, LPDDR5X and other advanced memory products.

An earlier estimate placed the cost of a complete Rubin NVL72 rack at approximately $9.1 million. That figure was higher than a previous estimate of $7.8 million because newer calculations included increased memory prices.

HBM4 alone could reportedly reach $53 per GB in 2027. At the scale used by a full AI rack, even a small increase in price per gigabyte can add hundreds of thousands of dollars to the finished system.

The reported changes would not reduce the 20.7TB of HBM4 attached to the GPUs. This memory is essential for handling large AI models and maintaining high accelerator performance.

Instead, NVIDIA appears to be targeting the LPDDR5X based SOCAMM capacity connected to the CPUs, where reductions may have a smaller effect on the rack’s main AI compute performance.

Cutting SOCAMM capacity could save hundreds of thousands of dollars

The reported financial impact depends on how far NVIDIA reduces the memory configuration.

If the LPDDR5X capacity is cut by half, the estimated cost could fall from around $1.2 million to approximately $586,000 per rack.

A more aggressive configuration using one quarter of the earlier memory capacity could reduce the estimated LPDDR5X cost to about $293,000.

These figures suggest that memory changes alone could remove hundreds of thousands of dollars from the cost of each Vera Rubin system.

The total rack price would not fall by the same percentage because GPUs, CPUs, networking, cooling, power equipment and other components remain expensive.

However, reducing the memory bill could help NVIDIA and its partners keep pricing closer to earlier expectations while protecting margins.

Large cloud providers and AI companies may also prefer lower memory configurations when their workloads do not require the maximum available CPU memory.

Vera Rubin still targets a major performance increase over Blackwell

The reported memory cuts do not appear to change NVIDIA’s main performance claims for the Rubin platform.

Early results shared for mixture of experts workloads suggested that a Vera Rubin NVL72 system could process around 800,000 tokens per second per megawatt.

A comparable Blackwell based platform was reported at roughly 80,000 tokens per second per megawatt, giving Rubin a claimed tenfold improvement under the measured workload.

These figures apply to a specific AI test and should not be treated as a universal tenfold increase across every application.

Real performance will depend on model size, precision, networking, software optimisation and how effectively each workload uses the available GPUs.

The unchanged HBM4 capacity suggests NVIDIA wants to preserve the accelerator memory needed for these demanding workloads while reducing memory elsewhere in the rack.

Long term agreements may give NVIDIA more flexibility than competitors

NVIDIA reportedly signed long term memory supply agreements before the current shortage became more severe.

Those arrangements may give the company better access to HBM and other memory products than smaller hardware suppliers. They may also provide more predictable pricing, although they cannot completely protect NVIDIA from wider market increases.

The decision to lower memory capacity suggests that supply availability alone is not the only concern. Even when components can be secured, their price can make the final system too expensive for some customers.

AI infrastructure buyers already face high spending on electricity, cooling, networking and data centre construction. A rack costing several million dollars must provide enough performance and utilisation to justify those additional expenses.

Offering different memory configurations could therefore make Vera Rubin accessible to a broader range of customers.

Final specifications have not been confirmed

The reported changes are based on an industry analysis rather than a complete public specification from NVIDIA.

The final Vera Rubin NVL72 configuration could vary by customer, workload or delivery date. Cloud providers may also request custom systems with more memory than the standard design.

It is not yet clear whether every rack will ship with 96GB SOCAMM modules or whether NVIDIA will offer several capacity options.

The company may also adjust pricing and specifications again if LPDDR5X or HBM4 supply improves before large volume shipments begin.

For now, the report shows how strongly memory prices are influencing the design of next generation AI systems. NVIDIA appears unwilling to reduce the GPU memory that directly supports model processing, but it may accept lower CPU memory capacity to prevent memory from consuming almost one third of the rack’s hardware cost.

Discover: News

Discussion (0)

Be the first to comment.