Analysis
OpenAI cut combined pricing on its GPT-5.6 Luna model by 80%, to $1.40 per million tokens, undercutting Google's cheapest Gemini tiers -- a move that landed one day after DeepSeek shipped a competitively priced V4 Flash refresh, making it as much a defensive response to pressure from both a US and a Chinese rival as an aggressive land-grab. In the same stretch, a very different part of the AI supply chain was moving in the opposite direction on price.
AI server demand is straining capacity well beyond TSMC's most advanced nodes, with mature-process foundries and advanced packaging providers now raising prices as supply constraints spread across the broader chip supply chain. Designers are increasingly holding unfulfilled orders because packaging and mature-node fabrication, not leading-edge wafer starts, have become the binding bottleneck. Apple CEO Tim Cook flagged the downstream version of the same problem on his final earnings call, warning of a coming 'hundred-year flood' in memory-chip pricing as AI data-center demand absorbs DRAM and HBM capacity that would otherwise go to consumer devices.
“Designers are increasingly holding unfulfilled orders because packaging and mature-node fabrication, not leading-edge wafer starts, have become the binding bottleneck.”
Put the two trends side by side and the margin math gets uncomfortable for frontier labs: API pricing is falling under competitive pressure from open-weight rivals, while the underlying compute and memory that inference actually runs on is getting more expensive, not less, because AI demand itself is straining the physical supply chain. That's a genuinely different dynamic than the software industry's usual cost curve, where falling prices to customers have historically tracked falling underlying infrastructure costs.
For investors in AI infrastructure and model companies, the practical question is which labs can absorb margin compression from both directions long enough to hold pricing power, and which get squeezed into raising prices again or slowing the pace of future cuts. Labs with their own chip supply commitments or long-term memory contracts locked in ahead of this squeeze are in a materially better position than those buying capacity on the spot market. What to watch: whether OpenAI, Google or DeepSeek adjust pricing again as memory and packaging costs continue climbing through the rest of the year.