Illustration for: DeepSeek's V4-Flash Is Now the Cheapest AI Model to Run

DeepSeek's V4-Flash Is Now the Cheapest AI Model to Run

DeepSeek's V4-Flash costs roughly 3 cents per benchmark test to run -- more than 100 times cheaper than Anthropic's Claude Fable 5 -- while scoring meaningfully lower on Artificial Analysis's Intelligence Index.

By the Numbers

~3 cents
V4-Flash cost/test
$3.15
Claude Fable 5 cost/test
50/100
V4-Flash score
9+ points
Gap to Opus/GPT-5.6
TC
By the AI Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
1 min read
ShareXLinkedInEmail
TC

The VC Read · Trace's Take

Trace Cohen

The bifurcation is the real story for any AI-native portfolio company: check whether your inference costs are actually workload-matched, because running everything on frontier-tier models when a nine-point capability gap doesn't matter for the task is pure margin erosion. Watch whether the cheap-tier models' capability gap narrows next -- that's what would actually threaten frontier-model pricing power, not another cost-per-token headline.

Analysis

DeepSeek's V4-Flash model costs an average of roughly 3 cents per benchmark test to run -- more than 100 times cheaper than Anthropic's Claude Fable 5, which averages $3.15 per test -- according to Artificial Analysis benchmark data reported by qz and Business Standard. On raw token pricing, V4-Flash runs $0.14 per million input tokens and $0.28 per million output tokens.

The catch is capability, not just cost. Artificial Analysis's Intelligence Index -- a composite of nine benchmarks spanning coding, reasoning and practical workplace tasks -- places V4-Flash at 50 out of 100, tying Google's Gemini 3.6 Flash and sitting one point behind Meta's Muse Spark 1.1 and Z.AI's GLM-5.2. Anthropic's Claude Opus 5 and Fable 5, along with OpenAI's GPT-5.6, all outscore V4-Flash by at least nine points on the same index.

Anthropic's Claude Opus 5 and Fable 5, along with OpenAI's GPT-5.6, all outscore V4-Flash by at least nine points on the same index.

Artificial Analysis's own framing matters here: it weights cost-per-test over sticker price specifically because the two can diverge sharply. A model that looks cheap per token can still produce an expensive bill if it needs many more reasoning steps to land on a correct answer -- meaning V4-Flash's headline price advantage doesn't automatically translate into a proportional real-world cost advantage for every workload, particularly ones where accuracy failures require retries.

For anyone building on top of frontier models, the practical read is a bifurcating market: V4-Flash and similarly-priced mid-tier models are becoming viable defaults for high-volume, lower-stakes tasks where a nine-point capability gap doesn't matter, while Opus-and-GPT-5.6-tier models hold their premium for workloads where the accuracy gap is the whole point. The interesting number to watch next isn't another price cut -- it's whether the capability gap between cheap and frontier models widens or narrows as each side iterates.

ShareXLinkedInEmail

Key Sources

3 sources

Reported by qz · First reported by Business Standard · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.