Analysis
Thinking Machines released Inkling-Small, a 276-billion-parameter open-weight model using just 12 billion active parameters per token -- roughly a quarter of the 41 billion its larger Inkling predecessor uses -- yet scoring 40 on the Artificial Analysis Intelligence Index against Inkling's 41, near-parity at a fraction of the compute cost per token. It actually outperforms its larger sibling on specific benchmarks including SWE-bench Verified. The release arrived just two weeks after Inkling itself shipped.
MiniMax followed a similar playbook in a different modality, open-sourcing H3, a multimodal video generation model capable of 2K resolution clips with native stereo sound, at under a third the cost per second of mainstream competitors -- with full weights released within days of the announcement, positioned deliberately against closed rivals including ByteDance's Seedance.
What connects the two releases isn't the modality, it's the cadence: both labs shipped a genuinely competitive smaller, cheaper sibling within roughly two weeks of their flagship model, not months or a year later as distillation work has traditionally trailed a primary release. That compression suggests distillation has become a standard part of the release pipeline itself, not a separate research track that ships whenever it happens to be ready.
The strategic logic is consistent across both labs: give away a smaller model that's genuinely close to flagship quality, and win developer mindshare and self-hosting distribution rather than only serving the high end of the cost-capability curve with one expensive, closed option. For investors, the pattern raises the bar for any startup positioning itself as 'the cheap, efficient alternative' to a big lab's flagship -- if labs themselves are shipping a competitive cheap tier within two weeks of their own flagship, that positioning window is closing fast.
What to watch: whether this small-sibling cadence becomes the default pattern across more labs with each future model generation, and how quickly enterprise adoption shifts toward the smaller, cheaper models once fine-tuning tooling catches up to match flagship-model workflows.