Illustration for: DeepSeek's New Flash Model Beats Opus 5 On Coding Tests

DeepSeek's New Flash Model Beats Opus 5 On Coding Tests

DeepSeek released V4.1-Flash, an open-weight model that scored 90.6 on Terminal-Bench 2.1 and resolved 74.2% of tasks on the DeepSWE coding benchmark, edging out Claude Opus 5 and GPT-5.6 Sol on both while pricing well below either.

By the Numbers

Sep 10, 2026
Released
90.6 (vs Opus 5: 89.1)
Terminal-Bench 2.1
74.2% (vs Opus 5: 74%)
DeepSWE v1.1
$0.15/M tokens
Input price (uncached)
552B (MoE)
Parameters
TC
By the AI Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
2 min read
ShareXLinkedInEmail
TC

The VC Read · Trace's Take

Trace Cohen

Any portfolio company running meaningful coding-agent spend on Claude or GPT should already be running a side-by-side eval against V4.1-Flash on their own workload -- the price gap is large enough that even a modest quality tradeoff can net out favorably, and open-weight self-hosting solves data-residency problems a US API can't. I'd discount the benchmark wins until someone outside DeepSeek reproduces them independently; self-reported scores from any lab, Chinese or American, deserve the same skepticism before they change a procurement decision.

Analysis

DeepSeek released V4.1-Flash on September 10, an open-weight, mixture-of-experts model with 552 billion parameters -- nearly double the 284 billion in its predecessor, V4-Flash -- that the company says outperforms its own larger flagship V4-Pro on several benchmarks, SiliconANGLE reported.

The published benchmark results are the real story: V4.1-Flash scored 90.6 on Terminal-Bench 2.1, narrowly ahead of Claude Opus 5's 89.1 and GPT-5.6 Sol's 88.8. On the DeepSWE v1.1 software-engineering benchmark, it resolved 74.2% of tasks, edging past Opus 5's 74% and well ahead of V4-Pro's 62.7%. Other reported scores include 88.1 on CyberGym, 65.4 on NL2Repo-Bench, 90.9 on GPQA Diamond and a Codeforces rating of 3,471. Against its own immediate predecessor, the jump is sharp: Terminal-Bench rose from 82.7 to 90.6 and DeepSWE from 54.4 to 74.2 in a single release cycle.

Pricing is where the gap against Western frontier labs is starkest. DeepSeek's Flash tier runs $0.003 per million cache-hit input tokens, $0.15 per million uncached input tokens, and $0.60 per million output tokens off-peak, with peak pricing double that -- a fraction of what Anthropic and OpenAI charge for models posting comparable or slightly lower benchmark scores. V4.1-Flash also ships as a genuinely open-weight model with a 1.0 million-token context window and multimodal input, meaning developers can download and run it on their own infrastructure rather than depending on an API.

The published benchmark results are the real story: V4.1-Flash scored 90.6 on Terminal-Bench 2.1, narrowly ahead of Claude Opus 5's 89.1 and GPT-5.6 Sol's 88.8.

DeepSeek besting Claude Opus 5 and GPT-5.6 on any published benchmark is notable given the compute and chip-access constraints Chinese AI labs operate under relative to their US counterparts -- export controls on advanced Nvidia GPUs into China have been a persistent headwind, and yet DeepSeek's architecture and training efficiency have repeatedly let it compete on raw benchmark performance despite that disadvantage, a pattern that has held across several of its recent releases.

Benchmark leadership on a handful of published tests is not the same as leadership across the much broader range of real-world enterprise and consumer tasks frontier models are actually used for -- coding and agentic benchmarks like Terminal-Bench and DeepSWE are narrower, more gameable proxies than the full range of reasoning, instruction-following and safety behavior that determines whether a model is trustworthy for production deployment, and neither Anthropic nor OpenAI has independently verified DeepSeek's self-reported scores.

For any company already running Claude or GPT-class models in coding-heavy workflows, V4.1-Flash's price-to-benchmark ratio is now aggressive enough that it's a legitimate cost-optimization evaluation, not just a curiosity -- open-weight availability also means it can be self-hosted for workloads with data-residency requirements that rule out a US-based API entirely.

Whether V4.1-Flash's benchmark scores hold up under independent, third-party evaluation rather than DeepSeek's own reporting, and how quickly Anthropic or OpenAI respond with a comparably priced tier, are the concrete signals that will determine whether this release meaningfully shifts developer workloads or remains a strong showing that doesn't move market share.

ShareXLinkedInEmail

Key Sources

2 sources

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.