2026-08-18-Tue · DeepSeek

From Issue 17 (2026-08-18) · 14 stories in this issue

❯ DeepSeek API Adopts Peak/Off-Peak Pricing; V4 Pro Peak Output Rises to 27 Yuan per Million Tokens

PRICINGStarting at 00:00 on August 17, DeepSeek is rolling out peak/off-peak time-of-use pricing for the DeepSeek API: 9:00–12:00 and 14:00–18:00 daily are peak windows, with the rest off-peak, and off-peak rates set at half the peak price. V4 Pro peak output is 27 yuan per million tokens, off-peak 13.5 yuan; V4 Flash comes in at 9 yuan and 4.5 yuan, respectively.

HIKESThis is not merely a discount arrangement; it is a genuine price increase. For V4 Pro, off-peak output pricing is up roughly 125% from the earlier initial pricing, and peak pricing is up roughly 350%; cache-hit input posts the steepest peak-time jump, at around 1100%. The official line is that pricing leverage will steer enterprise developers toward staggered scheduling, easing compute congestion and improving platform stability.

INFERENCE ECONThis marks the first time time-of-use electricity pricing has been ported into a large-model API — a sign that inference-side compute strain has pressed down into the pricing layer. What’s genuinely being rewritten is how batch workloads are scheduled — offline evaluation, bulk cleaning, overnight number-crunching, the kind of work that isn’t time-sensitive, now has a concrete cost rationale for shifting to off-peak windows. Rivals are watching too; once time-of-use proves effective, it’s only a matter of time before other vendors follow.

▪ SIGNALWhen APIs start billing by peak and off-peak windows, compute officially becomes a utility.