Illustration for: Claude Fable 5.1 Cuts Agentic Costs 45%

Claude Fable 5.1 Cuts Agentic Costs 45%

Anthropic launched Claude Fable 5.1 and Mythos 5.1, cutting cache-read pricing 75% to $0.25 per million tokens -- a change that can lower highly agentic workload costs by up to roughly 45% without touching headline per-token pricing.

By the Numbers

$0.25/M tokens
Cache read price (new)
75%
Cut vs. Fable 5
Up to ~45%
Agentic workload savings
$10 in / $50 out
Fable 5.1 pricing
52.6%
Terminal-Bench-Science
TC
By the AI Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
3 min read
ShareXLinkedInEmail
TC

The VC Read · Trace's Take

Trace Cohen

Cutting cache-read pricing instead of headline rates is Anthropic optimizing for the metric enterprise buyers actually track -- cost per completed agentic task, not cost per token -- and it's a smarter defensive move than a blanket price cut because it targets the exact workload where switching to a competitor is easiest. If you're underwriting any AI-native company's gross margins right now, model your agentic-workload cost assumptions off this 45% ceiling, not the flat 25%, because that's where the real competitive pressure is concentrated.

Analysis

Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 this week, with the headline change centered on cache-read pricing rather than a new headline per-token rate. Cache reads -- inputs the model has already processed and stored from earlier in a session -- now cost $0.25 per million tokens, a 75% reduction from Fable 5, VentureBeat reported. For typical workloads that cuts total cost by an estimated 25%; for highly agentic workloads -- coding agents, research agents, anything that repeatedly re-reads large amounts of cached context across a long session -- the savings can reach roughly 45%.

Claude Fable 5.1 is generally available at $10 per million input tokens and $50 per million output tokens. Claude Mythos 5.1 is described by Anthropic as the same underlying model with a different set of safeguards, available only through the company's trusted access programs designed specifically to support cybersecurity and life-sciences work -- a similar restricted-access structure to the one OpenAI is now using for Astra's advanced cyber capabilities. On Terminal-Bench-Science, a benchmark measuring scientific and technical agentic task performance, the new release scored 52.6%, per independent testing cited by MarkTechPost.

Why cache pricing, not headline pricing

The choice to cut cache-read costs rather than headline input/output pricing is a signal about where Anthropic sees its actual competitive pressure. Pulse covered Anthropic making its Sonnet 5 introductory pricing permanent rather than raising it as planned -- a defensive move against multi-model routing infrastructure that lets developers shift traffic to a cheaper competitor within days of a price change. Targeting cache-read costs specifically hits agentic and coding-assistant use cases hardest, precisely the workloads where developers run the same context through a model repeatedly and where competitors like OpenAI's GPT line and Google's Gemini models compete most directly on effective cost per completed task rather than sticker price per token.

The competitive landscape

Anthropic, OpenAI and Google have all been cutting or restructuring pricing this year rather than raising it, even as compute costs for training frontier models keep climbing -- a dynamic that shows up starkly in Anthropic's own roughly $90 billion of compute commitments signed in the past month. That combination -- rising infrastructure costs, falling or restructured customer-facing prices -- only works if inference costs are falling faster than headline pricing suggests, which is plausible given continued efficiency gains in serving infrastructure, but it does mean margin per query is under real pressure across the industry, not just at Anthropic.

What to watch next is whether OpenAI or Google respond with their own cache-pricing cuts specifically, rather than headline price moves -- that would confirm agentic workload cost is becoming the primary battleground for enterprise AI spend, not simple per-token pricing.

What the benchmark score adds

The 52.6% Terminal-Bench-Science result matters less as an absolute number than as a data point in a fast-moving benchmark race where labs leapfrog each other every few weeks on scientific and technical reasoning tasks. Combined with the pricing changes, it signals Anthropic is trying to win on both axes simultaneously -- capability and cost per completed task -- rather than picking one lane, a strategy that only works if the underlying model improvements are real rather than benchmark-optimized specifically for the headline number.

The Mythos 5.1 restricted-access structure is also worth noting on its own: pairing a capability upgrade with a narrower distribution model for the highest-risk use cases mirrors exactly the caution OpenAI is now applying to Astra, suggesting gated release for dual-use capabilities -- rather than broad simultaneous availability -- is becoming standard practice across frontier labs rather than a one-off policy choice at either company.

Additional reporting: TechCrunch.

ShareXLinkedInEmail

Key Sources

2 sources

Reported by VentureBeat · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.