DeepSeek V4-Pro is the focus of today’s briefing, with a practical twist: the model’s API lineup now uses peak and off-peak pricing, and that pricing regime became effective on August 16. 1
DeepSeek says V4-Pro is available in Expert Mode and through the API, adds adjustable reasoning effort, and supports the OpenAI Responses API with a Codex-oriented setup. The pricing documentation lists both V4-Flash and V4-Pro with OpenAI-compatible and Anthropic-compatible endpoints, one-million-token context, tool calls, JSON output, and Responses API support. 12
The operational change is the clock: DeepSeek says off-peak rates are half of peak rates. For V4-Pro, off-peak cache-hit input is $0.022 per million tokens, cache-miss input is $0.66, and output is $1.98; peak pricing doubles those figures. The page lists peak windows and concurrency limits for both models. 2
For developers, that points to a testable strategy: reserve V4-Pro and higher reasoning effort for difficult agent steps, use V4-Flash for routine volume, and schedule batch work around the published rate table. That is an operational inference from the pricing—not a claim that the cheaper model is equivalent.
One important boundary: DeepSeek’s GA announcement is dated August 13, so the model announcement should not be presented as a same-day release. The same announcement and DeepSeek’s official account identify August 16 as the date the new pricing takes effect. 13
DeepSeek describes production gains, but does not provide an independent benchmark in the announcement. Treat V4-Pro as a candidate for testing, then validate tool calls, recovery behavior, context-heavy tasks, cache hit rates, and spend before moving production traffic. 1