VC
Value Add VC
⚡HomePulse⚡Helpful Apps📝Blog🤝Partner
Illustration for: DeepSeek Raises V4 Prices Hours After Topping Agent Tests
Value Add VC/Pulse/AIDEEP DIVE$0.44/$1.32 per 1M tokens

DeepSeek Raises V4 Prices Hours After Topping Agent Tests

DeepSeek moved its V4 models to peak and off-peak pricing on August 16, raising rates across every tier, days after V4-Flash beat the company's own flagship on nine agent benchmarks at three cents per task.

By the Numbers

$0.44 / $1.32
V4-Flash peak input / output
$0.22 / $0.66
V4-Flash off-peak
~$0.03
Cost per agent task
9
Agent benchmarks Flash won
500 / 2,500
Concurrency: Pro / Flash
TC
By the AI Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
August 16, 2026
2 min read
ShareXLinkedInEmail
TC

The VC Read · Trace's Take

Trace Cohen

The cheapest model in the market just raised prices, and that is a bigger signal than any benchmark on the page. Two years of AI cost curves have been modeled by founders as monotonically down. Rebuild your unit economics with a flat inference cost and see whether the business still works. If your gross margin depends on DeepSeek staying at last month's rate, you do not have a margin, you have a subsidy.

AI Spending →

Analysis

DeepSeek switched its V4 API to time-of-day pricing at 16:00 UTC on August 16. V4-Flash now bills $0.44 per million input tokens and $1.32 per million output during peak windows -- 01:00 to 04:00 and 06:00 to 10:00 UTC -- and $0.22 in, $0.66 out at all other hours. As VentureBeat noted, the off-peak rate is still higher than the rate it replaced in every input, cache-hit and output category. This is a price increase wearing a discount's clothing.

Why the Timing Matters

The timing is what makes it interesting. On August 13, DeepSeek published a comparison putting V4-Flash-0731 against its own V4-Pro-0813, Claude Opus 4.8 and Fable 5 across Terminal Bench 2.1 and eight other agent benchmarks. Flash beat DeepSeek's own flagship on nine of them, at roughly $0.03 per task. V4-Pro went generally available the same day with DSpark speculative decoding, low/high/max reasoning-effort levels, a native OpenAI Responses API with one-click Codex setup, and published concurrency limits of 500 on Pro and 2,500 on Flash.

“Flash beat DeepSeek's own flagship on nine of them, at roughly $0.03 per task.”

The gap between benchmark performance and real agent work is the part builders should read twice. Terminal Bench scores measure task completion in constrained environments; production agents fail on long-horizon state, tool errors and ambiguous instructions that benchmarks do not model. A model that tops nine leaderboards and still stumbles on real workflows is the normal case, not an anomaly.

DeepSeek, founded in 2023 as a spinout of the quant fund High-Flyer, has been the primary downward force on frontier pricing since V3 landed in December 2024. That pressure has been reciprocated: Pulse previously covered OpenAI and Anthropic cutting prices specifically in response to Chinese labs. A price increase from DeepSeek is therefore a meaningful inflection -- it suggests either serving costs that finally had to be passed through, or capacity constraints that make demand-shaping more valuable than share gain.

For anyone running agents at volume, the operational answer is unglamorous: batch non-interactive workloads into the off-peak window and pin your model version. The 2,500-concurrency ceiling on Flash is the harder constraint for agent swarms, and it is the number that will bind before price does.

Context on how far the price floor has fallen: when DeepSeek's V3 landed in December 2024, frontier-class inference from U.S. labs ran roughly an order of magnitude higher per million tokens than what Flash charges even at its new peak rate. The compression since then has been the single largest input cost change AI application companies have experienced, and most business models built in 2025 quietly assumed it would continue.

The competitive set now prices in tiers rather than a single number. OpenAI and Anthropic have both cut list prices this year while pushing customers toward cached-input and batch endpoints; Google's Gemini Flash line competes directly on the cheap-and-fast axis; and open-weight deployments through providers like Together and Fireworks give buyers a floor they control themselves. DeepSeek's peak/off-peak move is the first time a major provider has used time-of-day as a demand-shaping lever, which is standard practice in electricity markets and new here.

ShareXLinkedInEmail

More on

DeepSeek →

Reported by VentureBeat · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.

Read Next

AI· Aug 17, 2026

AI Chip Stocks Rally as Anthropic's Blowout Quarter Lands

Illustration for: AI Chip Stocks Rally as Anthropic's Blowout Quarter Lands
AI

AI Chip Stocks Rally as Anthropic's Blowout Quarter Lands

Micron and Sandisk led premarket gains Monday after Anthropic's Q2 revenue surge bolstered the market's view that AI infrastructure spending will keep climbing rather than plateau.

AI· Aug 16, 2026

ChatGPT Can Now Log Every Click and Keystroke on Your Mac

Illustration for: ChatGPT Can Now Log Every Click and Keystroke on Your Mac
AI

ChatGPT Can Now Log Every Click and Keystroke on Your Mac

OpenAI's new Computer History feature records mouse clicks, typing and app switches on macOS to build a searchable timeline ChatGPT can reference, stored locally as unencrypted plain text and off by default.

AI· Aug 15, 2026

Eval Harness Finds Models Most Confident When Wrong

Illustration for: Eval Harness Finds Models Most Confident When Wrong
AI

Eval Harness Finds Models Most Confident When Wrong

A systematic evaluation harness surfaced what qualitative review missed: model confidence rises on the outputs that turn out to be wrong, inverting the signal teams rely on for routing and human escalation.

@Trace_Cohen·t@nyvp.com