Analysis
DeepSeek's V4-Flash model costs an average of roughly 3 cents per benchmark test to run -- more than 100 times cheaper than Anthropic's Claude Fable 5, which averages $3.15 per test -- according to Artificial Analysis benchmark data reported by qz and Business Standard. On raw token pricing, V4-Flash runs $0.14 per million input tokens and $0.28 per million output tokens.
The catch is capability, not just cost. Artificial Analysis's Intelligence Index -- a composite of nine benchmarks spanning coding, reasoning and practical workplace tasks -- places V4-Flash at 50 out of 100, tying Google's Gemini 3.6 Flash and sitting one point behind Meta's Muse Spark 1.1 and Z.AI's GLM-5.2. Anthropic's Claude Opus 5 and Fable 5, along with OpenAI's GPT-5.6, all outscore V4-Flash by at least nine points on the same index.
“Anthropic's Claude Opus 5 and Fable 5, along with OpenAI's GPT-5.6, all outscore V4-Flash by at least nine points on the same index.”
Artificial Analysis's own framing matters here: it weights cost-per-test over sticker price specifically because the two can diverge sharply. A model that looks cheap per token can still produce an expensive bill if it needs many more reasoning steps to land on a correct answer -- meaning V4-Flash's headline price advantage doesn't automatically translate into a proportional real-world cost advantage for every workload, particularly ones where accuracy failures require retries.
For anyone building on top of frontier models, the practical read is a bifurcating market: V4-Flash and similarly-priced mid-tier models are becoming viable defaults for high-volume, lower-stakes tasks where a nine-point capability gap doesn't matter, while Opus-and-GPT-5.6-tier models hold their premium for workloads where the accuracy gap is the whole point. The interesting number to watch next isn't another price cut -- it's whether the capability gap between cheap and frontier models widens or narrows as each side iterates.