Analysis
Anthropic and OpenAI both released cheaper flagship-adjacent models within hours of each other this week: Claude Opus 5.5 costs 40% less to run than Opus 5, priced at $4 per million input tokens and $20 per million output tokens, while OpenAI's GPT-6 Sol charges $2/$10 per million tokens and its lighter GPT-6 Luna undercuts both at just $0.10/$0.50, according to CNBC's reporting on the releases.
xAI's Grok 4.7, released days earlier with a 500,000-token context window, and Alibaba's Qwen3.8-Omni-Flash, whose per-hour audio-input price Alibaba says is more than 98% lower than its own prior omnimodal model, round out a competitive field where every major lab cut prices within roughly the same two-week window rather than any single company moving alone.
The pattern across all of these cuts is consistent: labs are pricing a cheaper, faster tier for high-volume tasks (summarization, extraction, simple agent loops) well below their flagship reasoning-model pricing, while keeping premium pricing on the most capable models for complex coding and agentic work -- a two-tier structure that increasingly resembles cloud-compute pricing tiers more than a single per-model price point.
For founders building on these APIs, the practical read is that per-token cost is falling faster than most product margin models assumed a year ago, which is good news for gross margins on AI-native products but bad news for any startup whose moat was primarily 'we made inference cheaper' rather than a genuine product or data advantage -- that specific competitive angle is being commoditized by the labs themselves roughly every quarter now.