Analysis
OpenAI cut combined token pricing on its GPT-5.6 Luna model by 80% this week, bringing the blended input-output rate down to $1.40 per million tokens -- 20 cents for input, $1.20 for output -- from $7 previously. The mid-tier Terra model dropped a further 20% in the same announcement, OpenAI's steepest simultaneous price moves of the year.
The new Luna pricing directly undercuts Google's cheapest Gemini tiers: Gemini 3.5 Flash-Lite runs $2.80 combined and Gemini 3.6 Flash runs $9, meaning OpenAI is now the cheapest major-lab option for the high-throughput, low-latency workloads -- summarization, classification, routing -- where price sensitivity is highest and margins are thinnest.
“OpenAI's move reads as much as a defensive response to DeepSeek and Gemini's cheaper tiers as it does an aggressive land-grab.”
In the same announcement, OpenAI launched a premium "Sol Fast" mode for its flagship GPT-5.6 Sol model, priced at $70 per million tokens combined for 2.5x throughput over the standard tier. Cutting the cheap tier 80% while introducing a 5x-priced premium tier in the same breath effectively splits OpenAI's lineup into a race-to-the-bottom commodity layer and a high-margin premium layer simultaneously.
The timing adds pressure from multiple directions at once: the cuts land one day after DeepSeek released DeepSeek-V4-Flash-0731, a competitively priced Chinese model squeezing the same low-cost segment from the other side. OpenAI's move reads as much as a defensive response to DeepSeek and Gemini's cheaper tiers as it does an aggressive land-grab.
What to watch: whether Google or DeepSeek respond with further cuts of their own, and how quickly startups currently differentiating on "cheaper than GPT" pricing have to rework that pitch now that the price gap has narrowed sharply.