DeepSeek V4 Flash Undercuts OpenAI as Moonshot Expands With 20,000 NVIDIA H200 GPUs

news
DeepSeek V4 Flash Undercuts OpenAI as Moonshot Expands With 20,000 NVIDIA H200 GPUs

DeepSeek has introduced an updated version of its V4 Flash model with aggressive API pricing, adding pressure to an emerging price competition between Chinese and US artificial intelligence companies.

The refreshed model, called DeepSeek V4 Flash 0731, is priced at $0.14 per million input tokens and $0.28 per million output tokens. Those rates are substantially lower than the recently reduced pricing announced for OpenAI’s GPT-5.6 Luna.

Moonshot is also expanding its computing resources. The company has reportedly secured access to a cluster containing 20,000 NVIDIA H200 GPUs through Alibaba, giving it significantly more capacity to train and improve future versions of its Kimi models.

Together, the developments show that Chinese AI laboratories are competing through both lower inference prices and larger training infrastructure.

DeepSeek V4 Flash Targets Low Cost Inference

DeepSeek V4 Flash 0731 reportedly contains 284 billion parameters and is positioned as a fast, lower cost model.

The model is said to deliver performance comparable with Anthropic’s Opus 4.8 in some evaluations, despite claims that the latter may use a much larger parameter count. Parameter estimates for closed models remain uncertain, so direct comparisons should be treated carefully.

The clearest difference is pricing.

ModelInput price per million tokensOutput price per million tokens
DeepSeek V4 Flash 0731$0.14$0.28
OpenAI GPT-5.6 Luna after discount$0.20$1.20
OpenAI GPT-5.6 Luna before discount$1.00$6.00

DeepSeek’s output pricing is less than one quarter of OpenAI’s newly discounted rate. Input pricing is also lower, although the gap is smaller.

For developers running large workloads, output tokens often account for a substantial part of the total bill. That makes the $0.28 rate particularly relevant for applications involving long responses, document generation, coding, and automated research.

Low pricing alone does not determine value. Reliability, latency, context limits, tool use, safety controls, uptime, and regional availability also affect whether a model is suitable for production.

OpenAI Had Just Reduced Its Own Prices

OpenAI recently cut GPT-5.6 Luna pricing by as much as 80 percent.

Input tokens fell from $1 per million to $0.20, while output tokens dropped from $6 to $1.20. The company attributed the reduction to architectural and efficiency improvements.

The move was widely viewed as an effort to strengthen OpenAI’s position against lower cost competitors.

DeepSeek’s response arrived only hours later and removed much of that pricing advantage.

Pricing changePrevious rateNew rateReduction
GPT-5.6 Luna input$1.00$0.2080%
GPT-5.6 Luna output$6.00$1.2080%
DeepSeek V4 Flash inputNot stated$0.14Not stated
DeepSeek V4 Flash outputNot stated$0.28Not stated

The speed of the response suggests that pricing has become a central competitive tool. AI companies are trying to reduce inference costs while maintaining model quality.

This can benefit developers, but it may also place pressure on smaller providers that cannot operate at the same scale.

Moonshot Adds a Large H200 Cluster

Moonshot has reportedly obtained access to 20,000 NVIDIA H200 GPUs through Alibaba.

The H200 is designed for large AI training and inference workloads. It includes high bandwidth memory and is widely used for frontier model development.

A cluster of that size could allow Moonshot to train larger models, run more experiments, and shorten development cycles.

Moonshot expansionReported detail
GPU modelNVIDIA H200
GPU count20,000
Infrastructure providerAlibaba
Main useAI model training and inference
Likely beneficiaryFuture Kimi models

The arrangement also illustrates how Chinese AI developers are obtaining computing resources through domestic cloud and technology providers.

Access to advanced accelerators remains one of the main constraints facing AI laboratories. Large models require substantial computing power, memory capacity, networking, and electricity.

Securing thousands of GPUs gives Moonshot more freedom to develop the next generation of its models without relying on small or fragmented clusters.

Claims Around Model Distillation Remain Disputed

Moonshot has faced allegations that its Kimi models benefited from distillation involving systems developed by Anthropic.

Distillation is a process in which one model learns from the outputs of another. It can be used legitimately in some settings, but it becomes controversial when the original provider has not authorised the use of its model outputs.

The reference also includes allegations that Moonshot accessed advanced NVIDIA systems through countries outside China, including Thailand.

These claims have not been proven in the material provided. They should therefore be treated as allegations rather than established facts.

The wider dispute reflects growing tension between the United States and China over access to advanced processors, model training methods, and intellectual property.

Thinking Machines Introduces Inkling-Small

Thinking Machines has also released an open-weight model called Inkling-Small.

The model reportedly contains 276 billion parameters and supports native reasoning across audio and images. It also allows variable reasoning effort, which may help developers balance cost and performance.

On the Artificial Analysis Intelligence Index cited in the report, DeepSeek V4 Flash 0731 received a score of 50 percent. Inkling-Small scored 40 percent, matching the earlier DeepSeek V4 Flash version.

ModelReported parametersIntelligence Index score
DeepSeek V4 Flash 0731284 billion50%
Inkling-Small276 billion40%
Earlier DeepSeek V4 FlashNot stated40%

Benchmark scores can provide a useful comparison, but they do not represent every real workload. Performance may vary across coding, mathematics, reasoning, image understanding, audio tasks, and long context use.

Open weights can also make a model more valuable for organisations that want to fine-tune or deploy it on their own infrastructure.

AI Competition Is Moving Beyond Model Quality

The current AI market is increasingly shaped by four factors: capability, price, openness, and access to computing hardware.

DeepSeek is competing strongly on price. Moonshot is increasing its training capacity. Thinking Machines is offering an open-weight alternative. OpenAI is reducing costs to defend its position.

Competitive areaExample
Lower inference costDeepSeek V4 Flash 0731
Larger training infrastructureMoonshot’s 20,000 H200 GPUs
Open-weight availabilityInkling-Small
Major price reductionsGPT-5.6 Luna
Multimodal reasoningInkling-Small audio and image support

For developers, this could lead to lower API bills and more model choices. It may also make switching between providers easier when several systems offer similar performance.

The reported pricing gives DeepSeek a clear cost advantage on paper, particularly for output tokens. Whether that advantage holds in production will depend on model reliability, speed, access restrictions, and actual task quality.

The larger trend is already visible. AI companies are no longer competing only to produce the strongest model. They are also competing to make advanced intelligence cheaper to run and easier to scale.

Discover: News

Discussion (0)

Be the first to comment.