2026-08-01-Sat · DeepSeek · OpenAI

From Issue 1 (2026-08-01) · 12 stories in this issue

❯ DeepSeek Ships V4-Flash Stable: Weights Open-Sourced Same Day, Agent Benchmarks Overtake Its Own Pro Preview

OPEN-SOURCEDeepSeek released the V4-Flash-0731 stable build on July 31. The architecture hasn’t changed one bit; the weights went up on a public Hugging Face repo under the MIT license the same day, and the API entered public beta in lockstep. Across the nine agent and coding benchmarks the company published, this compact model — 284B total parameters, 13B activated — beat its own larger V4-Pro preview across the board. The technical report simply carries over the V4 paper from April, making the point explicit: the model hasn’t changed; the training has.

GAINSThe improvement is concentrated in agent tasks. April’s preview had long drawn criticism in this category, and official data shows DeepSWE jumping from 7.3 to 54.4 — nearly sevenfold — while Terminal Bench 2.1 rose from 61.8 to 82.7, leaving the V4-Pro preview’s 72.1 behind. Third-party numbers line up: Artificial Analysis hands it an Intelligence Index of 50, tying Google’s Gemini 3.6 Flash and a full 10 points above April’s preview. And that 10-point edge comes entirely from post-training — total parameters, activated parameters, and the 1-million-token context window are all untouched, and pricing hasn’t moved either. DeepSeek also noted that this upgrade applies only to the V4-Flash endpoint; the V4-Pro API and web portal stay unchanged for now, with the Pro stable release coming “as soon as possible.” The new endpoint natively supports the Responses format and is Codex-compatible.

COSTThe move stings most for the closed-source models stuck in the middle tier. From above, OpenAI has just cut GPT-5.6 Luna’s input price by 80%; from below, an open-source model with comparable intelligence and give-away weights has appeared — the mid-range price band is squeezed from both ends. For Chinese teams building agent products, the inference-cost line item can essentially be crossed out of critical decisions; the gating factors become evaluation ecosystems and engineering reliability. Overseas closed-source vendors, meanwhile, face a harder question: when the same architecture gains 10 points purely from post-training, how much premium does a model generation still command?

▪ SIGNALRun the same model twice, and the second pass is worth 10 extra points — post-training is becoming a more valuable asset than parameter count.