Analysis
DeepSeek released V4.1-Flash on September 10, an open-weight, mixture-of-experts model with 552 billion parameters -- nearly double the 284 billion in its predecessor, V4-Flash -- that the company says outperforms its own larger flagship V4-Pro on several benchmarks, SiliconANGLE reported.
The published benchmark results are the real story: V4.1-Flash scored 90.6 on Terminal-Bench 2.1, narrowly ahead of Claude Opus 5's 89.1 and GPT-5.6 Sol's 88.8. On the DeepSWE v1.1 software-engineering benchmark, it resolved 74.2% of tasks, edging past Opus 5's 74% and well ahead of V4-Pro's 62.7%. Other reported scores include 88.1 on CyberGym, 65.4 on NL2Repo-Bench, 90.9 on GPQA Diamond and a Codeforces rating of 3,471. Against its own immediate predecessor, the jump is sharp: Terminal-Bench rose from 82.7 to 90.6 and DeepSWE from 54.4 to 74.2 in a single release cycle.
Pricing is where the gap against Western frontier labs is starkest. DeepSeek's Flash tier runs $0.003 per million cache-hit input tokens, $0.15 per million uncached input tokens, and $0.60 per million output tokens off-peak, with peak pricing double that -- a fraction of what Anthropic and OpenAI charge for models posting comparable or slightly lower benchmark scores. V4.1-Flash also ships as a genuinely open-weight model with a 1.0 million-token context window and multimodal input, meaning developers can download and run it on their own infrastructure rather than depending on an API.
“The published benchmark results are the real story: V4.1-Flash scored 90.6 on Terminal-Bench 2.1, narrowly ahead of Claude Opus 5's 89.1 and GPT-5.6 Sol's 88.8.”
DeepSeek besting Claude Opus 5 and GPT-5.6 on any published benchmark is notable given the compute and chip-access constraints Chinese AI labs operate under relative to their US counterparts -- export controls on advanced Nvidia GPUs into China have been a persistent headwind, and yet DeepSeek's architecture and training efficiency have repeatedly let it compete on raw benchmark performance despite that disadvantage, a pattern that has held across several of its recent releases.
Benchmark leadership on a handful of published tests is not the same as leadership across the much broader range of real-world enterprise and consumer tasks frontier models are actually used for -- coding and agentic benchmarks like Terminal-Bench and DeepSWE are narrower, more gameable proxies than the full range of reasoning, instruction-following and safety behavior that determines whether a model is trustworthy for production deployment, and neither Anthropic nor OpenAI has independently verified DeepSeek's self-reported scores.
For any company already running Claude or GPT-class models in coding-heavy workflows, V4.1-Flash's price-to-benchmark ratio is now aggressive enough that it's a legitimate cost-optimization evaluation, not just a curiosity -- open-weight availability also means it can be self-hosted for workloads with data-residency requirements that rule out a US-based API entirely.
Whether V4.1-Flash's benchmark scores hold up under independent, third-party evaluation rather than DeepSeek's own reporting, and how quickly Anthropic or OpenAI respond with a comparably priced tier, are the concrete signals that will determine whether this release meaningfully shifts developer workloads or remains a strong showing that doesn't move market share.