2026-08-03-Mon · DeepSeek

From Issue 3 (2026-08-03) · 10 stories in this issue

❯ DeepSeek V4-Flash official release goes live, output price cut to 2 yuan per million tokens — one-twelfth of Pro

PRICE SLASHDeepSeek’s V4-Flash official release entered public beta on July 31, with output priced at 2 yuan per million tokens and input at 1 yuan. Against sibling flagship V4-Pro’s 24 yuan output / 12 yuan input, Flash’s output price comes in at one-twelfth that of Pro — while outscoring the Pro preview on agentic benchmarks. Overseas researcher teortaxes puts it even more bluntly: 14x cheaper than V4-Pro, 50% faster, and better to use.

NO CAPABILITY CUTThe key shift: this price cut comes with no capability discount. On agentic tests, V4-Flash official release scored 25.2 points, closing in on Claude Opus 4.8’s 25.7 points, while the V4-Pro preview managed only 15.8 — the first time a budget tier has run flush against the top closed-source line. Overseas developer communities report from hands-on testing that V4-Flash’s tool-calling framework is far better than Kimi K3. Analyst kimmonismus posted a Pareto-frontier chart, marking it the current price-performance winner. teortaxes argues that among Chinese labs, only Luna is currently competing in the same market.

BUDGETS REDRAWNThe first to redraw the math are teams building agent products. Splitting traffic — cheap models on simple tasks, premium models on complex chains — was a compromise forced by pricing. When a budget tier scores 25, the split itself becomes redundant engineering cost. Pressure will hit first on token-billing intermediary providers — once inference costs fall into this range, the spread from reselling call credits is essentially erased.

▪ SIGNALThe moment a cheap model catches up to a premium one, what gets eliminated isn’t the premium model — it’s model routing as a business.