The AI Inference Price War, By The Numbers logo

The AI Inference Price War, By The Numbers

Anthropic and OpenAI cut inference prices within hours of each other this week, and a look at the current per-token pricing across the major labs shows just how far the cost of frontier intelligence has fallen in 2026.

By the Numbers

$4/$20 per 1M
Opus 5.5
$2/$10 per 1M
GPT-6 Sol
$0.10/$0.50 per 1M
GPT-6 Luna
-40% vs Opus 5
Opus 5.5 cost cut
TC
By the AI Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
1 min read
ShareXLinkedInEmail

THE RUNDOWN

1

Every major AI lab -- Anthropic, OpenAI, xAI, Alibaba -- cut inference prices within roughly the same two-week window, evidence of a genuine industry-wide price war rather than one company's isolated move.

2

The emerging two-tier pricing structure, cheap high-volume models alongside expensive flagship reasoning models, increasingly resembles cloud-compute pricing tiers more than a single per-model price point.

3

Falling per-token costs improve gross margins for AI-native products built on these APIs faster than most margin models assumed a year ago.

4

Startups whose competitive moat was primarily cheaper inference face commoditization roughly every quarter now, as the labs themselves cut prices faster than most startups can differentiate on cost alone.

TC

The VC Read · Trace's Take

Trace Cohen

The two-tier pricing structure -- cheap high-volume model, expensive flagship reasoning model -- is the real story here, not any single price point, because it means labs have converged on cloud-style tiered pricing rather than one number per model. For anyone underwriting an AI-native startup's margins, model the flagship-tier price into your worst case and the cheap-tier price into your best case, because which tier a given workload actually needs is usually the biggest swing factor in unit economics.

Analysis

Anthropic and OpenAI both released cheaper flagship-adjacent models within hours of each other this week: Claude Opus 5.5 costs 40% less to run than Opus 5, priced at $4 per million input tokens and $20 per million output tokens, while OpenAI's GPT-6 Sol charges $2/$10 per million tokens and its lighter GPT-6 Luna undercuts both at just $0.10/$0.50, according to CNBC's reporting on the releases.

xAI's Grok 4.7, released days earlier with a 500,000-token context window, and Alibaba's Qwen3.8-Omni-Flash, whose per-hour audio-input price Alibaba says is more than 98% lower than its own prior omnimodal model, round out a competitive field where every major lab cut prices within roughly the same two-week window rather than any single company moving alone.

The pattern across all of these cuts is consistent: labs are pricing a cheaper, faster tier for high-volume tasks (summarization, extraction, simple agent loops) well below their flagship reasoning-model pricing, while keeping premium pricing on the most capable models for complex coding and agentic work -- a two-tier structure that increasingly resembles cloud-compute pricing tiers more than a single per-model price point.

For founders building on these APIs, the practical read is that per-token cost is falling faster than most product margin models assumed a year ago, which is good news for gross margins on AI-native products but bad news for any startup whose moat was primarily 'we made inference cheaper' rather than a genuine product or data advantage -- that specific competitive angle is being commoditized by the labs themselves roughly every quarter now.

ShareXLinkedInEmail

Key Sources

3 sources

Reported by CNBC · First reported by Value Add Pulse · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.