VC
Value Add VC
⚡HomePulse⚡Helpful Apps📝Blog🤝Partner
Illustration for: Small Open Models Are Becoming Labs' New Funnel Strategy
Value Add VC/Pulse/AI

Small Open Models Are Becoming Labs' New Funnel Strategy

Thinking Machines' Inkling-Small matches its flagship within one point at a quarter of the compute, and MiniMax open-sourced a competitive video model at a third of rivals' cost -- distillation is now a release strategy, not a side project.

By the Numbers

276B (12B active)
Inkling-Small params
40 vs 41
Score vs flagship
<1/3 per second
H3 cost vs rivals
~2 weeks
Gap between releases
TC
Trace Cohen
Early-stage VC & angel · Founder, New York Venture Partners
August 3, 2026
1 min read
ShareXLinkedInEmail

THE RUNDOWN

1

Thinking Machines released Inkling-Small, a 276-billion-parameter open-weight model using just 12 billion active parameters per token versus 41 billion for its larger Inkling predecessor, yet scoring 40 on the Artificial Analysis Intelligence Index against Inkling's 41 -- near-parity at roughly a quarter of the active compute

2

MiniMax open-sourced H3, a multimodal video generation model producing 2K clips with native stereo sound at under a third the cost per second of mainstream competitors, with full weights following within days

3

Both releases arrived within roughly two weeks of their respective flagship models, not a year later -- evidence that distillation has become a standard, fast-follow release cadence rather than a separate, slower research project

4

The strategic logic is the same in both cases: give away a genuinely competitive smaller model to win developer mindshare and distribution, rather than only serving the high end of the cost-capability curve with a single flagship

TC

The VC Read · Trace's Take

Trace Cohen

A quarter of the active compute for one point of benchmark difference, shipped two weeks after the flagship rather than a year later, should worry any startup whose entire pitch is 'we're the cheap, efficient alternative to a big lab's model' -- that positioning window just got a lot shorter, because the labs are now doing it themselves as a standard part of the release cadence.

AI Valuations Tracker →

Analysis

Thinking Machines released Inkling-Small, a 276-billion-parameter open-weight model using just 12 billion active parameters per token -- roughly a quarter of the 41 billion its larger Inkling predecessor uses -- yet scoring 40 on the Artificial Analysis Intelligence Index against Inkling's 41, near-parity at a fraction of the compute cost per token. It actually outperforms its larger sibling on specific benchmarks including SWE-bench Verified. The release arrived just two weeks after Inkling itself shipped.

MiniMax followed a similar playbook in a different modality, open-sourcing H3, a multimodal video generation model capable of 2K resolution clips with native stereo sound, at under a third the cost per second of mainstream competitors -- with full weights released within days of the announcement, positioned deliberately against closed rivals including ByteDance's Seedance.

What connects the two releases isn't the modality, it's the cadence: both labs shipped a genuinely competitive smaller, cheaper sibling within roughly two weeks of their flagship model, not months or a year later as distillation work has traditionally trailed a primary release. That compression suggests distillation has become a standard part of the release pipeline itself, not a separate research track that ships whenever it happens to be ready.

The strategic logic is consistent across both labs: give away a smaller model that's genuinely close to flagship quality, and win developer mindshare and self-hosting distribution rather than only serving the high end of the cost-capability curve with one expensive, closed option. For investors, the pattern raises the bar for any startup positioning itself as 'the cheap, efficient alternative' to a big lab's flagship -- if labs themselves are shipping a competitive cheap tier within two weeks of their own flagship, that positioning window is closing fast.

What to watch: whether this small-sibling cadence becomes the default pattern across more labs with each future model generation, and how quickly enterprise adoption shifts toward the smaller, cheaper models once fine-tuning tooling catches up to match flagship-model workflows.

ShareXLinkedInEmail

More on

Thinking Machines →MiniMax →

Reported by Value Add Pulse Analysis · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.

Read Next

AI· Aug 6, 2026

IonQ Lands $28M DARPA Deal for Atomic Clocks

Illustration for: IonQ Lands $28M DARPA Deal for Atomic Clocks
AI$28M contract

IonQ Lands $28M DARPA Deal for Atomic Clocks

IonQ secured a $28 million DARPA contract extension to scale production of its Evergreen-05 optical atomic clocks, expanding beyond quantum computing into defense-grade timing hardware.

AI· Aug 7, 2026

Why the AI Labs Just Rewired Their Org Charts

Illustration for: Why the AI Labs Just Rewired Their Org Charts
AI

Why the AI Labs Just Rewired Their Org Charts

Hassabis moving to chair, Jeff Dean's exit, and Anthropic's new chip team all landed in one week -- a trace take on what it means that frontier labs are restructuring around infrastructure, not research.

AI· Aug 7, 2026

Palantir Jumps 10% as BofA Turns Bullish

Illustration for: Palantir Jumps 10% as BofA Turns Bullish
AI+10.3% stock move

Palantir Jumps 10% as BofA Turns Bullish

Palantir shares rose 10.3% after Bank of America issued a bullish note following the company's blowout Q2 earnings -- even as BofA's own market-wide sentiment gauge flashes a rare sell signal.

@Trace_Cohen·t@nyvp.com