❯ DeepSeek Quietly Lists V4-Pro-0813 with 1M Context and 384K Max Output
LISTINGDeepSeek has added DeepSeek-V4-Pro-0813 to the “Models & Pricing” page of its official API documentation without a changelog. The flagship version offers a 1M token context window and a maximum output of 384,000 tokens, uses thinking mode by default, and includes a separate non-thinking-mode endpoint for latency-sensitive calls.
PRICINGAccording to the API docs, pricing remains at $0.435 per million input tokens, $0.87 for output, with cached input as low as $0.003625. The V4 Flash launched in July had already pulled this tier’s price down, and the flagship now follows the same standard. In the same tier, compared with Grok 4.6’s $2 input, this is more than four times cheaper. However, DeepSeek has warned users that prices may rise significantly later.
CADENCEPutting the price list up before any announcement is DeepSeek’s typical release style — letting developers stumble upon it in the API first, then letting community benchmarks complete the narrative. What to watch is that price-increase warning: whether this price is the new normal determines whether teams building long-context applications dare to include it in their cost models for the next year, and whether this price cut is a competitive strategy or a capacity dividend.
▪ SIGNALThe four-times-cheaper flagship model carries its own price-increase warning; cost models need to leave a line for it.