VC
Value Add VC
⚡HomePulse⚡Helpful Apps📝Blog🤝Partner
Illustration for: DeepSeek and Peking University Open-Source DSpark, Speeding LLM Inference by Up to 85%
Value Add VC/Pulse/AIUp to 85% faster

DeepSeek and Peking University Open-Source DSpark, Speeding LLM Inference by Up to 85%

DeepSeek and Peking University open-sourced DSpark, a speculative-decoding framework that boosts per-user LLM generation speed by 60-85% -- and up to 661% throughput under tight latency constraints -- without hardware upgrades or model retraining. Released under the MIT license and already live in DeepSeek's V4-Flash and V4-Pro production models, it is another Chinese efficiency breakthrough aimed at slashing the cost of serving AI.

By the Numbers

60-85% per user
Speedup
Up to 661%
Throughput Gain
MIT
License
DeepSeek V4-Flash / V4-Pro
Live On
Peking University
Co-developer
TC
By the AI Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
June 29, 2026
2 min read
ShareXLinkedInEmail

THE RUNDOWN

1

Cutting inference cost by software alone undercuts the need for ever-more expensive accelerators

2

It is a second major Chinese AI release in 48 hours, alongside Meituan's LongCat-2.0

3

MIT-licensed and model-agnostic (supports Qwen, Gemma), it can spread across the ecosystem

4

Efficiency, not just scale, is becoming China's counter to the US chip embargo

TC

The VC Read · Trace's Take

Trace Cohen

Pair DSpark with Meituan's LongCat the same week and you see China's actual strategy under the chip embargo: if you can't buy more compute, extract more from what you have. An 85% inference speedup given away for free is a direct shot at the 'just buy more GPUs' reflex -- and because DeepSpec is model-agnostic, the gains leak into Qwen, Gemma and the whole open ecosystem. For founders, cheaper inference is pure margin upside, whoever ships it. The caveat is the usual one with vendor numbers: speculative decoding can trade accuracy for speed, so wait for independent benchmarks before you rip out your serving stack.

⚡ AI Chip Wars →🤖 AI Landscape →

Analysis

DeepSeek and Peking University have jointly open-sourced DSpark, a speculative-decoding framework that accelerates large language model inference by 60% to 85% per user -- and delivers up to a 661% throughput gain under strict latency constraints -- with no hardware upgrades or model retraining required, according to VentureBeat. It is released under the MIT license on GitHub, alongside DeepSpec, a general-purpose codebase for training custom draft models.

The technical approach attacks a known limitation of speculative decoding. Classic methods train a separate, smaller 'draft' model to propose tokens that the larger target model then verifies -- effective, but costly to build and maintain. DSpark instead grafts the speculative head directly onto the target model, reducing layer duplication, and pairs a 'semi-autoregressive generation' method with a 'confidence-scheduled verification' system. The deployed configuration, 'DSpark-5,' improves per-user generation speed by 60-85% on DeepSeek-V4-Flash and 57-78% on V4-Pro.

“The deployed configuration, 'DSpark-5,' improves per-user generation speed by 60-85% on DeepSeek-V4-Flash and 57-78% on V4-Pro.”

The significance is economic. Inference -- the perpetual cost of serving a model with every query -- is the dominant and fastest-growing line item in AI, the same dynamic driving billion-dollar bets on Baseten, Groq and Upscale AI. A free, open framework that wrings 60-85% more speed out of existing hardware attacks that cost from the software side, reducing the pressure to buy ever-more accelerators. For a Chinese ecosystem constrained by US export controls on the most advanced GPUs, squeezing more out of available compute is a strategic necessity, not just an optimization.

The timing makes a pattern. DSpark landed within 48 hours of Meituan open-sourcing the 1.6-trillion-parameter LongCat-2.0 -- two MIT-licensed releases from Chinese players in the same window, both aimed at efficiency and openness. Crucially, DeepSpec is model-agnostic, with configurations supporting Alibaba's Qwen and Google's Gemma, so DSpark's gains can spread well beyond DeepSeek's own models and into the broader open-source community.

The bear case is that vendor-reported speedups need independent verification, real-world gains vary by workload, and speculative decoding can trade accuracy for speed if poorly tuned. Western enterprises may also hesitate to build inference infrastructure around Chinese-origin frameworks regardless of license. What to watch: independent benchmarks of DSpark across model families, how quickly the open-source community adopts DeepSpec, and whether efficiency breakthroughs like this meaningfully blunt the impact of US chip restrictions.

ShareXLinkedInEmail

More on

DeepSeek →

Reported by VentureBeat · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.

Read Next

AI· Aug 14, 2026

OpenAI Sheds Senior Execs in Pre-IPO Shakeup

Illustration for: OpenAI Sheds Senior Execs in Pre-IPO Shakeup
AI

OpenAI Sheds Senior Execs in Pre-IPO Shakeup

OpenAI has lost its chief revenue officer, its longtime COO and several senior leaders within days of each other, as co-founder Greg Brockman consolidates operating control ahead of a planned public listing.

AI· Aug 13, 2026

Anthropic's CFO Starts Courting IPO Investors

Illustration for: Anthropic's CFO Starts Courting IPO Investors
AI

Anthropic's CFO Starts Courting IPO Investors

Anthropic CFO Krishna Rao has begun early, informal meetings with prospective IPO investors, though he has not discussed valuation -- the $2 trillion figure circulating on Wall Street comes from investors' own math, not from Anthropic.

AI· Aug 13, 2026

Gemini 3.7 Flash Launches With 50% Price Cut for Coding

Illustration for: Gemini 3.7 Flash Launches With 50% Price Cut for Coding
AI

Gemini 3.7 Flash Launches With 50% Price Cut for Coding

Google released Gemini 3.7 Flash just three weeks after 3.6 Flash, cutting introductory API pricing in half while improving coding, debugging and enterprise-automation benchmarks over its predecessor.

@Trace_Cohen·t@nyvp.com