Illustration for: Sakana AI Splits Fugu Into Max And Ultra v2

Sakana AI Splits Fugu Into Max And Ultra v2

Sakana AI split its multi-agent orchestration model into a cost-first Fugu Max, priced 40-60% below Sonnet 5 and GPT 5.6 Terra, and a capability-first Fugu Ultra v2 that leads on two benchmarks.

By the Numbers

$2/$6 per 1M tok
Fugu Max pricing
40-60% cheaper
Price vs rivals
48.3 vs Opus 27.3
Ultra v2, Chartography
74.3
Ultra v2, DeepSWE
TC
By the AI Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
2 min read
ShareXLinkedInEmail
TC

The VC Read · Trace's Take

Trace Cohen

The real signal isn't the benchmark scores, it's that a mid-tier lab can now credibly split a single model family into two price points and still claim frontier-adjacent performance on the expensive tier. What I'd diligence for any portfolio company with meaningful inference spend: whether their current model-selection logic can absorb a 40-60% price swing on their primary vendor without a re-architecture, because that's now happening roughly every quarter.

Analysis

Sakana AI launched two new models, Fugu Max and Fugu Ultra v2, splitting its orchestration system into separate cost-first and capability-first tiers, MarkTechPost reported. Fugu Max prices at $2 per million input tokens and $6 per million output tokens -- 40-60% below Sonnet 5, GPT 5.6 Terra, and Kimi K3 -- while Fugu Ultra v2 is tuned purely for answer quality on difficult, multi-step problems, scoring 48.3 on the Chartography benchmark against Opus 5's 27.3, and 74.3 on DeepSWE.

Both models ship through the same OpenAI-compatible API endpoint Sakana AI introduced with the original Fugu launch in June, letting existing users switch tiers with what the company describes as a one-line code change. That original release was a multi-agent orchestration system that routed tasks across a swappable pool of frontier LLMs rather than a single model competing head-to-head with GPT or Claude; splitting it into Max and Ultra v2 is Sakana AI's answer to a problem every orchestration layer eventually hits -- customers want either the cheapest reliable option or the best possible one, and a single middle-of-the-road tier satisfies neither.

The pricing move lands squarely inside what CNBC recently termed "model fatigue" -- labs shipping version after version as the differentiation between them compresses.

The pricing move lands squarely inside what CNBC recently termed "model fatigue" -- labs shipping version after version as the differentiation between them compresses. Sakana AI's answer to that fatigue isn't a bigger model, it's a cheaper one: Fugu Max competes on unit economics for the high-volume coding-assistant and chatbot workloads that make up most inference spend, while Fugu Ultra v2 chases the smaller slice of customers who'll pay a premium for genuinely harder multi-step reasoning.

For developers building on top of any of these models, the risk isn't picking the wrong tier today -- it's that frontier labs are re-pricing and re-splitting their lineups fast enough that a cost model built around today's rates can be stale within a quarter. Sakana AI is a much smaller lab than OpenAI, Anthropic, or Google, and its ability to compete on price this aggressively suggests inference costs across the industry are falling faster than most application-layer startups have priced into their own margins.

ShareXLinkedInEmail

Key Sources

3 sources

Reported by MarkTechPost · First reported by AlphaSignal · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.