2026-08-05-Wed · DeepSeek

From Issue 5 (2026-08-05) · 9 stories in this issue

❯ DeepSeek’s Updated V4-Flash Matches GLM-5.2 in Benchmarks at One-Tenth the Price

10X PRICE GAPDeepSeek’s updated V4-Flash pushes the price-performance curve for open-source models down another notch: researcher Nathan Lambert says the new version now matches GLM-5.2 in benchmarks, while OpenRouter’s listing reportedly shows V4-Flash at $0.14 per million input tokens and $0.28 per million output, versus $1.40 and $4.40 for GLM-5.2. That works out to a 10x gap on input and 15x on output. It is currently the No. 1 model on OpenRouter by call volume.

UNDERESTIMATED ADOPTIONAnother remark from Lambert flags an easily missed detail: adoption of the original V4-Flash has been severely underestimated in discussions — real usage is far higher than the outside impression, and related activity on HuggingFace is just as strong. The ecosystem has been quick to follow — third parties have already released 14 quantized versions, from lossless BF16 all the way down to 1-bit, all in GGUF format for direct loading by local inference runtimes. Worth calling out separately is how this batch is evaluated: no benchmark comparisons — instead, KL divergence measures how far the compressed output distribution drifts from the original weights, asking “is this model still the original model,” not “how many points can it still score.”

BUDGETS REWRITTENFor teams building products, token cost is no longer the main budget line on the application side — based on the public pricing above, what really eats money at this tier is context management and call counts. What gets squeezed are closed-source APIs sitting in the middle tier: above them, frontier models command a capability premium; below them, open weights match benchmarks at a tenth of the price — the middle layer has a hard time explaining what it charges for. 1-bit quantization plus GGUF also pulls in another group — those who can run near-frontier models on a personal machine, a cohort that was never in the pricing table’s consideration set before.

▪ SIGNALWhen benchmark parity costs a tenth of the price, the middle-tier API has no story left to tell.