Companies
OpenAI Cuts GPT-5.6 Luna 80% as DeepSeek Closes the Gap
OpenAI cut GPT-5.6 Luna prices by 80 percent on July 30. A day later, DeepSeek's retrained V4 Flash landed one point behind it on benchmarks at roughly 60 percent lower cost per task.

When a company slashes a price by 80 percent overnight, is that confidence or pressure? OpenAI would like Thursday's cut to read as an engineering story. The calendar tells a second one.
As of July 30, GPT-5.6 Luna, the smallest model in OpenAI's lineup, costs $0.20 per million input tokens and $1.20 per million output tokens, both down 80 percent. The official explanation is efficiency: GPT-5.6 Sol reportedly optimized OpenAI's own GPU software, trimming deployment costs by about 20 percent, while speculative decoding lifted token generation speed by more than 15 percent. The engineering is probably real. The timing, with low-cost Chinese providers squeezing the volume segment, is hard to ignore.
DeepSeek answered within a day
On July 31, DeepSeek released V4 Flash "0731," a retrained version of its budget model. It scores 50 on the Artificial Analysis Intelligence Index, ten points above its predecessor and a single point behind Luna. On GDPval, a benchmark built around real office work, it jumped from 1,189 to 1,559 Elo. It also completes tasks using about 12 percent fewer tokens than before.
The pricing comparison is the uncomfortable part for OpenAI: per task, DeepSeek still comes in roughly 60 percent cheaper than Luna even after the cut, helped by a 98 percent cache discount that beats the industry's usual 90. Cache discounts matter more than they sound: the fixed instructions and documents you resend with every request are processed almost free from the second call onward, and in chatbot and agent workloads that repetition is most of the bill. The weights are on Hugging Face under an MIT license: 284 billion total parameters, 13 billion active, and a one-million-token context window.
So who wins a price war?
Buyers do, at least in the short run. High-volume, low-glamour workloads such as classification, summarization, and support bots now cost a fraction of what they did a year ago, which quietly rewrites the business case for automation projects that once failed on cost alone. The catch: any price/performance analysis older than a few months is obsolete, and this week it aged in 24 hours. Keep your stack model-agnostic so a provider swap is a configuration change, revisit pricing whenever a major release lands, and stay a little skeptical of pure efficiency framings. In a market this competitive, nobody gives up 80 percent of a price for fun; make sure quality and rate limits are pinned down in writing before you commit volume to the cheapest tier.
Sources: The Decoder (OpenAI), The Decoder (DeepSeek)

Written by
Muhammet Fatih Batman
Founder & Editor
Founder of YZ Uzman, with 20+ years of experience in web design and software development.