Analysis
Thinking Machines released Inkling-Small, a 276-billion-parameter open-weight multimodal reasoning model, just two weeks after debuting its larger Inkling model. Despite using only 12 billion active parameters per token compared with Inkling's 41 billion, Inkling-Small scores 40 on the Artificial Analysis Intelligence Index against Inkling's 41 -- near performance parity at a fraction of the active compute cost per token.
The model accepts text, image and audio inputs and produces text output, with a context window scaling up to 1 million tokens. On specific benchmarks, Inkling-Small actually outperforms its larger sibling -- 80.2% versus 77.6% on SWE-bench Verified, and 64.7% versus 63.8% on Terminal-Bench 2.1 -- a reminder that parameter count and active compute don't map linearly to every task's performance ceiling.
“The model accepts text, image and audio inputs and produces text output, with a context window scaling up to 1 million tokens.”
Thinking Machines released full weights under an Apache 2.0 license on Hugging Face, with fine-tuning support through its Tinker API, and is offering a limited-time 50% discount on standard-context API pricing at launch. The strategy pairs a frontier-scale flagship model with a genuinely competitive, cheaper sibling rather than serving only the high end of the cost-capability curve -- a deliberate bet against the one-size-fits-all foundation model approach that dominated the industry's first several years.
For investors and technical buyers, Inkling-Small is a data point in the broader trend of model providers -- both open and closed -- discovering that smaller, well-distilled models can match much larger ones on many real-world tasks, undercutting the assumption that frontier capability requires frontier parameter counts. What to watch: adoption of Inkling-Small relative to Inkling itself once enterprise fine-tuning data comes in through Tinker, and whether Thinking Machines continues shipping a small/large pairing with each future model generation.