VC
Value Add VC
⚡HomePulse⚡Helpful Apps📝Blog🤝Partner
Illustration for: Small Open Models Are Becoming Labs' New Funnel Strategy
Value Add VC/Pulse/AI

Small Open Models Are Becoming Labs' New Funnel Strategy

Thinking Machines' Inkling-Small matches its flagship within one point at a quarter of the compute, and MiniMax open-sourced a competitive video model at a third of rivals' cost -- distillation is now a release strategy, not a side project.

276B (12B active)
Inkling-Small params
40 vs 41
Score vs flagship
<1/3 per second
H3 cost vs rivals
~2 weeks
Gap between releases
TC
Trace Cohen
Early-stage VC & angel · Founder, New York Venture Partners
August 3, 2026
1 min read
ShareXLinkedInEmail

THE RUNDOWN

1

Thinking Machines released Inkling-Small, a 276-billion-parameter open-weight model using just 12 billion active parameters per token versus 41 billion for its larger Inkling predecessor, yet scoring 40 on the Artificial Analysis Intelligence Index against Inkling's 41 -- near-parity at roughly a quarter of the active compute

2

MiniMax open-sourced H3, a multimodal video generation model producing 2K clips with native stereo sound at under a third the cost per second of mainstream competitors, with full weights following within days

3

Both releases arrived within roughly two weeks of their respective flagship models, not a year later -- evidence that distillation has become a standard, fast-follow release cadence rather than a separate, slower research project

4

The strategic logic is the same in both cases: give away a genuinely competitive smaller model to win developer mindshare and distribution, rather than only serving the high end of the cost-capability curve with a single flagship

TC

The VC Read · Trace's Take

Trace Cohen

A quarter of the active compute for one point of benchmark difference, shipped two weeks after the flagship rather than a year later, should worry any startup whose entire pitch is 'we're the cheap, efficient alternative to a big lab's model' -- that positioning window just got a lot shorter, because the labs are now doing it themselves as a standard part of the release cadence.

AI Valuations Tracker →

Analysis

Thinking Machines released Inkling-Small, a 276-billion-parameter open-weight model using just 12 billion active parameters per token -- roughly a quarter of the 41 billion its larger Inkling predecessor uses -- yet scoring 40 on the Artificial Analysis Intelligence Index against Inkling's 41, near-parity at a fraction of the compute cost per token. It actually outperforms its larger sibling on specific benchmarks including SWE-bench Verified. The release arrived just two weeks after Inkling itself shipped.

MiniMax followed a similar playbook in a different modality, open-sourcing H3, a multimodal video generation model capable of 2K resolution clips with native stereo sound, at under a third the cost per second of mainstream competitors -- with full weights released within days of the announcement, positioned deliberately against closed rivals including ByteDance's Seedance.

What connects the two releases isn't the modality, it's the cadence: both labs shipped a genuinely competitive smaller, cheaper sibling within roughly two weeks of their flagship model, not months or a year later as distillation work has traditionally trailed a primary release. That compression suggests distillation has become a standard part of the release pipeline itself, not a separate research track that ships whenever it happens to be ready.

The strategic logic is consistent across both labs: give away a smaller model that's genuinely close to flagship quality, and win developer mindshare and self-hosting distribution rather than only serving the high end of the cost-capability curve with one expensive, closed option. For investors, the pattern raises the bar for any startup positioning itself as 'the cheap, efficient alternative' to a big lab's flagship -- if labs themselves are shipping a competitive cheap tier within two weeks of their own flagship, that positioning window is closing fast.

What to watch: whether this small-sibling cadence becomes the default pattern across more labs with each future model generation, and how quickly enterprise adoption shifts toward the smaller, cheaper models once fine-tuning tooling catches up to match flagship-model workflows.

ShareXLinkedInEmail

Analysis and editorial commentary by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.

Read Next

AI· Aug 1, 2026

OpenAI's Astra Cracks 10 Unsolved Math Problems

Illustration for: OpenAI's Astra Cracks 10 Unsolved Math Problems
AI

OpenAI's Astra Cracks 10 Unsolved Math Problems

An internal version of OpenAI's next model, Astra, solved ten previously-open math and computer-science problems for roughly $2,000 in compute -- proofs Fields Medalist Timothy Gowers says he'd back for a top journal.

AI· Aug 3, 2026

AI's Dual-Use Risk Just Became a Live Security Problem

Illustration for: AI's Dual-Use Risk Just Became a Live Security Problem
AI

AI's Dual-Use Risk Just Became a Live Security Problem

A scoping error that let Claude models breach real production systems, and a US-China robot-ban standoff threatening rare-earth retaliation, broke days apart -- proof AI's dual-use risk is now operational, not hypothetical.

AI· Aug 3, 2026

DeepSeek's Offensive Use Forces a Red-Team Rethink

Illustration for: DeepSeek's Offensive Use Forces a Red-Team Rethink
AI

DeepSeek's Offensive Use Forces a Red-Team Rethink

Palo Alto Networks caught an operator using DeepSeek to autonomously attack 460+ systems after Claude and OpenAI's models refused the same job -- a live test of whether model-level refusal is a real safety layer or just a routing problem.

@Trace_Cohen·t@nyvp.com