Analysis
OpenAI cut API prices for two models in its GPT-5.6 lineup on Thursday, reducing Luna's cost by 80% to $0.20 per million input tokens and $1.20 per million output tokens, down from $1 and $6, and cutting Terra's rate by 20% to $2 and $12 per million tokens, down from $2.50 and $15. The company also introduced a Fast mode for GPT-5.6 Sol, its higher-end model, offering up to 2.5x faster processing at twice the standard price, replacing the previous Priority Processing tier.
OpenAI attributed the price cuts to efficiency gains made during GPT-5.6's development, including the model's own ability to rewrite and optimize its production inference code and improvements in speculative decoding that speed up token generation. Notably, the lower Luna and Terra prices flow through automatically to how usage is counted in Codex and ChatGPT Work, meaning existing subscribers effectively get more usage for the same spend without any action required.
The cuts land amid an intensifying API pricing war across frontier labs, as Anthropic, Google, and OpenAI all compete for developer and enterprise usage on cost per token as much as raw capability. An 80% cut on a flagship-tier model is an unusually large single move, suggesting OpenAI's inference costs on Luna specifically had more room to compress than the company had previously passed through to customers.
For founders building on OpenAI's API, the price cuts materially change unit economics for high-volume, cost-sensitive workloads, particularly anything currently running on Luna. What to watch: whether Anthropic and Google respond with comparable cuts of their own, and whether OpenAI's claimed efficiency gains show up in its own margin structure when the company eventually discloses more detail ahead of its IPO.