Models
DeepSeek V4-Flash Outscores the Flagship at the Same Price
DeepSeek's official V4-Flash release keeps the architecture and the $0.14 price but redoes post-training, and now beats the company's own V4-Pro-Preview on all nine published agent benchmarks.

Fourteen cents per million input tokens. That price didn't move when DeepSeek shipped the official release of V4-Flash on July 31. What moved is everything else: the new build, designated V4-Flash-0731, beats DeepSeek's own higher-tier V4-Pro-Preview on all nine of the agent and coding benchmarks the company published. The budget model now outperforms the flagship it was meant to sit under.
Same architecture, rebuilt training
Technically, the release is an interesting case study in where model gains come from these days. V4-Flash-0731 keeps the exact architecture and size of the preview version: 284 billion total parameters in a mixture-of-experts design with 13 billion active per query, and a 1-million-token context window. DeepSeek didn't scale anything up. It redid only the post-training, focused on agentic behavior, tool use, and coding, and that alone transformed the model's character.
Two benchmark lines tell the story
On Terminal Bench 2.1, the new Flash scores 82.7, against 72.1 for V4-Pro-Preview and 61.8 for the old Flash preview. On DeepSWE, which measures real software engineering tasks, the jump is steeper still: from 7.3 in the preview to 54.4. The model also speaks the Responses API format natively and has been tuned to slot into Codex-style coding agents.
The usual caution applies: every number above comes from DeepSeek's own published evaluations, and independent testing hasn't landed yet. Recent launches from Kimi and Qwen showed how wide the gap between launch benchmarks and third-party results can get. Still, the underlying claim is technically plausible rather than pure marketing: post-training alone producing gains of this size, without touching the architecture, is a pattern other labs have demonstrated repeatedly this year.
Cheap models stopped being the compromise
This release is best read as the next move in the price war we covered days ago, when OpenAI cut GPT-5.6 Luna by 80 percent and DeepSeek answered within a day. Now DeepSeek answers again, this time with capability instead of price. For engineering teams and businesses running AI in production, the practical shift is that defaulting to the most expensive model is becoming a habit rather than a decision. In cost-sensitive workloads like coding assistance, agent automation, and bulk document processing, meaningfully better models keep arriving at the same low price point every few months. Teams that treat model selection as a quarterly review item, instead of an annual commitment, are the ones pocketing the difference.
Sources: DeepSeek, MarkTechPost

Written by
Muhammet Fatih Batman
Founder & Editor
Founder of YZ Uzman, with 20+ years of experience in web design and software development.