Analysis
Sakana AI launched two new models, Fugu Max and Fugu Ultra v2, splitting its orchestration system into separate cost-first and capability-first tiers, MarkTechPost reported. Fugu Max prices at $2 per million input tokens and $6 per million output tokens -- 40-60% below Sonnet 5, GPT 5.6 Terra, and Kimi K3 -- while Fugu Ultra v2 is tuned purely for answer quality on difficult, multi-step problems, scoring 48.3 on the Chartography benchmark against Opus 5's 27.3, and 74.3 on DeepSWE.
Both models ship through the same OpenAI-compatible API endpoint Sakana AI introduced with the original Fugu launch in June, letting existing users switch tiers with what the company describes as a one-line code change. That original release was a multi-agent orchestration system that routed tasks across a swappable pool of frontier LLMs rather than a single model competing head-to-head with GPT or Claude; splitting it into Max and Ultra v2 is Sakana AI's answer to a problem every orchestration layer eventually hits -- customers want either the cheapest reliable option or the best possible one, and a single middle-of-the-road tier satisfies neither.
“The pricing move lands squarely inside what CNBC recently termed "model fatigue" -- labs shipping version after version as the differentiation between them compresses.”
The pricing move lands squarely inside what CNBC recently termed "model fatigue" -- labs shipping version after version as the differentiation between them compresses. Sakana AI's answer to that fatigue isn't a bigger model, it's a cheaper one: Fugu Max competes on unit economics for the high-volume coding-assistant and chatbot workloads that make up most inference spend, while Fugu Ultra v2 chases the smaller slice of customers who'll pay a premium for genuinely harder multi-step reasoning.
For developers building on top of any of these models, the risk isn't picking the wrong tier today -- it's that frontier labs are re-pricing and re-splitting their lineups fast enough that a cost model built around today's rates can be stale within a quarter. Sakana AI is a much smaller lab than OpenAI, Anthropic, or Google, and its ability to compete on price this aggressively suggests inference costs across the industry are falling faster than most application-layer startups have priced into their own margins.