Analysis
DeepSeek switched its V4 API to time-of-day pricing at 16:00 UTC on August 16. V4-Flash now bills $0.44 per million input tokens and $1.32 per million output during peak windows -- 01:00 to 04:00 and 06:00 to 10:00 UTC -- and $0.22 in, $0.66 out at all other hours. As VentureBeat noted, the off-peak rate is still higher than the rate it replaced in every input, cache-hit and output category. This is a price increase wearing a discount's clothing.
Why the Timing Matters
The timing is what makes it interesting. On August 13, DeepSeek published a comparison putting V4-Flash-0731 against its own V4-Pro-0813, Claude Opus 4.8 and Fable 5 across Terminal Bench 2.1 and eight other agent benchmarks. Flash beat DeepSeek's own flagship on nine of them, at roughly $0.03 per task. V4-Pro went generally available the same day with DSpark speculative decoding, low/high/max reasoning-effort levels, a native OpenAI Responses API with one-click Codex setup, and published concurrency limits of 500 on Pro and 2,500 on Flash.
“Flash beat DeepSeek's own flagship on nine of them, at roughly $0.03 per task.”
The gap between benchmark performance and real agent work is the part builders should read twice. Terminal Bench scores measure task completion in constrained environments; production agents fail on long-horizon state, tool errors and ambiguous instructions that benchmarks do not model. A model that tops nine leaderboards and still stumbles on real workflows is the normal case, not an anomaly.
DeepSeek, founded in 2023 as a spinout of the quant fund High-Flyer, has been the primary downward force on frontier pricing since V3 landed in December 2024. That pressure has been reciprocated: Pulse previously covered OpenAI and Anthropic cutting prices specifically in response to Chinese labs. A price increase from DeepSeek is therefore a meaningful inflection -- it suggests either serving costs that finally had to be passed through, or capacity constraints that make demand-shaping more valuable than share gain.
For anyone running agents at volume, the operational answer is unglamorous: batch non-interactive workloads into the off-peak window and pin your model version. The 2,500-concurrency ceiling on Flash is the harder constraint for agent swarms, and it is the number that will bind before price does.
Context on how far the price floor has fallen: when DeepSeek's V3 landed in December 2024, frontier-class inference from U.S. labs ran roughly an order of magnitude higher per million tokens than what Flash charges even at its new peak rate. The compression since then has been the single largest input cost change AI application companies have experienced, and most business models built in 2025 quietly assumed it would continue.
The competitive set now prices in tiers rather than a single number. OpenAI and Anthropic have both cut list prices this year while pushing customers toward cached-input and batch endpoints; Google's Gemini Flash line competes directly on the cheap-and-fast axis; and open-weight deployments through providers like Together and Fireworks give buyers a floor they control themselves. DeepSeek's peak/off-peak move is the first time a major provider has used time-of-day as a demand-shaping lever, which is standard practice in electricity markets and new here.