NVIDIA faces new pressure as AI engineers weigh power and cooling costs

news
NVIDIA faces new pressure as AI engineers weigh power and cooling costs

NVIDIA’s AI chips still hold a strong position in the data center market, but Evercore ISI analysts say some AI engineers are paying closer attention to the full cost of using them. The concern is not only chip performance, but also power use, cooling needs, utilization, and cost per token.

NVIDIA has often argued that its AI platforms offer better total cost of ownership because they deliver strong performance per watt. Morgan Stanley recently said Blackwell based data centers may cost about twice as much to build compared with custom AI chip alternatives, but can deliver much higher performance per watt.

The shift from training to inference is changing how AI hardware is judged

Evercore says the AI market is moving from a training heavy phase toward an inference led phase. That changes buying priorities. Instead of only looking at maximum throughput and bandwidth, hyperscalers and AI engineers are now looking more closely at cost per token, return on investment, power, cooling, utilization, and overall ownership cost.

That shift could help custom ASICs and other accelerators gain more attention. Evercore says some engineers are willing to use ASICs or “good enough” alternatives if those options improve economics. The report also says NVIDIA’s claims around large performance gains are not always enough to convince engineers who are focused on operating costs.

AreaWhat engineers are watching
Cost per tokenHow much it costs to generate AI output
Power useElectricity needed to run the hardware
CoolingExtra infrastructure needed to manage heat
UtilizationHow much of the hardware is actually used
TCOFull cost of owning and running the system
AlternativesCustom ASICs and other accelerators

A Nebius expert cited in the report said inference demand makes up as much as 95 percent of enterprise workload use cases. That helps explain why buyers are looking at hardware differently. For many companies, the question is no longer only which chip is fastest, but which platform can produce useful AI output at the lowest practical cost.

This story is related to the earlier NVIDIA Vera Rubin and memory cost reports, but it focuses on a different pressure point. Those reports covered NVIDIA’s next AI platforms and rising component costs. This one shows why some hyperscalers may still look beyond NVIDIA, especially as inference workloads make power, cooling, and cost per token more important.

Discover: News

Discussion (0)

Be the first to comment.