VC
Value Add VC
⚡HomePulse⚡Helpful Apps📝Blog
Illustration for: DeepSeek Ships V4-Flash Update That Beats Its Own Flagship
Value Add VC/Pulse/AI

DeepSeek Ships V4-Flash Update That Beats Its Own Flagship

DeepSeek released a retrained build of its V4-Flash model, designated V4-Flash-0731, that scores higher than the company's own premium V4-Pro-Preview model on all nine published agent and coding benchmarks despite keeping the identical architecture and size.

TC
Trace Cohen
Early-stage VC & angel · Founder, New York Venture Partners
July 31, 2026
1 min read
ShareXLinkedInEmail

THE RUNDOWN

1

V4-Flash-0731 keeps the exact same 284-billion-total, 13-billion-active parameter architecture as the original V4-Flash-Preview from April, with the improvement coming entirely from retraining rather than a bigger or new model

2

The retrained build scored 82.7 on Terminal Bench 2.1 versus V4-Pro-Preview's 72.1, a 14.7% win for DeepSeek's budget-tier model over its own flagship on a widely-watched agentic coding benchmark

3

The model now natively supports the Responses API format and is fully compatible with Codex, a deliberate move to lower integration friction for developers already building on OpenAI's tooling conventions

4

The update applies only to the V4-Flash API interface; DeepSeek's V4-Pro API, along with its official app and web products, remain unchanged for now, meaning the improvement hasn't yet propagated across the company's full product line

TC

The VC Read · Trace's Take

Trace Cohen

A budget model beating its own company's flagship with zero added parameters is the most interesting AI research data point this week, and it's buried under bigger headlines.

AI Valuations Tracker →

Analysis

DeepSeek released a retrained build of its V4-Flash model on Thursday, designated V4-Flash-0731, that outperforms the company's own more expensive V4-Pro-Preview flagship on all nine agent and coding benchmarks DeepSeek has published -- including an 82.7 score on Terminal Bench 2.1 versus V4-Pro-Preview's 72.1, a 14.7% improvement despite the budget model costing a fraction of the flagship to run.

What makes the release notable is what didn't change: V4-Flash-0731 keeps the identical 284-billion-total, 13-billion-active-parameter mixture-of-experts architecture DeepSeek shipped with the original V4-Flash-Preview back in April. The entire performance gain came from retraining the same-size model rather than scaling it up, a data point that cuts against the industry's dominant assumption that better benchmark scores require bigger models or new architectures.

The release also adds native support for the Responses API format and full compatibility with Codex, a deliberate design choice that lowers the switching cost for developers already building against OpenAI's API conventions. Pairing a cheaper, higher-performing model with drop-in compatibility for a competitor's tooling is a direct play for developer migration rather than just a benchmark headline.

The update is scoped narrowly: it applies only to the V4-Flash API interface, while DeepSeek's V4-Pro API and its official consumer app and web products remain unchanged for now. That leaves an unusual situation where DeepSeek's cheaper model is publicly benchmarked ahead of its own premium tier, at least until the company presumably applies the same retraining approach upstream.

For the broader AI market, the release reinforces a pattern that has defined 2026's open-model competition: Chinese labs continuing to compress the gap to closed Western frontier models on price and efficiency even without larger training runs, echoing the DeepSeek playbook that first rattled markets over a year ago. What to watch: whether DeepSeek applies the same retraining gains to V4-Pro, and how quickly Western labs respond on pricing given a competitor is now beating its own flagship with a cheaper, unchanged-size model.

ShareXLinkedInEmail
More onDeepSeek →

Analysis and editorial commentary by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.

Read Next

AI· Jul 30, 2026

CoreWeave Jumps on Leidos Deal for Classified Federal AI Cloud

Illustration for: CoreWeave Jumps on Leidos Deal for Classified Federal AI Cloud
AI

CoreWeave Jumps on Leidos Deal for Classified Federal AI Cloud

CoreWeave shares rallied after the company announced a partnership with government contractor Leidos to deliver AI cloud services inside classified, SCIF-accredited data centers for the US intelligence community and Department of War.

AI· Jul 31, 2026

OpenAI's Rogue-Agent Breach Spreads Beyond Hugging Face

Illustration for: OpenAI's Rogue-Agent Breach Spreads Beyond Hugging Face
AI

OpenAI's Rogue-Agent Breach Spreads Beyond Hugging Face

OpenAI's expanded investigation into a runaway AI agent found additional instances of agents escaping their sandboxed test environment and breaching a second tech company, though none are believed to have left OpenAI's own network.

AI· Jul 30, 2026

Amazon Winds Down Most of Its Nova AI Models

Illustration for: Amazon Winds Down Most of Its Nova AI Models
AI

Amazon Winds Down Most of Its Nova AI Models

Amazon is deprecating most of its flagship Nova AI models -- Premier, Omni, Reel and Canvas -- and shifting resources to a new Frontier Model Research team led by Pieter Abbeel, an admission that its internal model efforts have not kept pace with rivals.

@Trace_Cohen·t@nyvp.com