VC
Value Add VC
⚡HomePulse⚡Helpful Apps📝Blog🤝Partner
Illustration for: OpenAI's Price Cuts Are Outrunning Its Own Chip Costs
Value Add VC/Pulse/AI

OpenAI's Price Cuts Are Outrunning Its Own Chip Costs

OpenAI cut GPT-5.6 Luna pricing 80% the same stretch that AI demand pushed chip packaging and mature-node manufacturing prices higher across the semiconductor supply chain -- a margin squeeze from both directions at once.

80%
Luna price cut
$1.40/1M tokens
New Luna price
Packaging, mature nodes
Chip bottleneck
'100-year flood'
Memory warning
TC
Trace Cohen
Early-stage VC & angel · Founder, New York Venture Partners
August 3, 2026
2 min read
ShareXLinkedInEmail

THE RUNDOWN

1

OpenAI cut combined GPT-5.6 Luna pricing 80% to $1.40 per million tokens, undercutting Google's cheapest Gemini tiers, one day after DeepSeek released a competitively priced V4 Flash refresh -- a defensive move against pressure from both a US and a Chinese rival at once

2

In the same stretch, AI server demand strained capacity across mature-node chip manufacturing and advanced packaging -- not just TSMC's most advanced nodes -- pushing prices higher as designers hold orders they can't get fabricated fast enough

3

That capacity strain connects to Apple CEO Tim Cook's warning on his final earnings call about a coming 'hundred-year flood' in memory pricing, as AI data-center demand absorbs DRAM and HBM capacity that would otherwise serve consumer devices

4

The result is a labs' margin structure being squeezed from both directions simultaneously: API prices falling under competitive pressure even as the underlying hardware and memory that inference actually runs on gets more expensive, not less

TC

The VC Read · Trace's Take

Trace Cohen

Cutting API prices 80% while the memory and packaging that inference runs on gets more expensive is a bet that competitive pressure matters more than near-term margin -- a bet only a handful of labs with the balance sheet to eat both sides of that squeeze can actually afford to make. Every startup pricing its own product against 'today's cheap API rate' should model what happens if that rate gets pulled back once the underlying hardware cost pressure catches up.

AI Valuations Tracker →OpenAI API Pricing in 2026 →

Analysis

OpenAI cut combined pricing on its GPT-5.6 Luna model by 80%, to $1.40 per million tokens, undercutting Google's cheapest Gemini tiers -- a move that landed one day after DeepSeek shipped a competitively priced V4 Flash refresh, making it as much a defensive response to pressure from both a US and a Chinese rival as an aggressive land-grab. In the same stretch, a very different part of the AI supply chain was moving in the opposite direction on price.

AI server demand is straining capacity well beyond TSMC's most advanced nodes, with mature-process foundries and advanced packaging providers now raising prices as supply constraints spread across the broader chip supply chain. Designers are increasingly holding unfulfilled orders because packaging and mature-node fabrication, not leading-edge wafer starts, have become the binding bottleneck. Apple CEO Tim Cook flagged the downstream version of the same problem on his final earnings call, warning of a coming 'hundred-year flood' in memory-chip pricing as AI data-center demand absorbs DRAM and HBM capacity that would otherwise go to consumer devices.

“Designers are increasingly holding unfulfilled orders because packaging and mature-node fabrication, not leading-edge wafer starts, have become the binding bottleneck.”

Put the two trends side by side and the margin math gets uncomfortable for frontier labs: API pricing is falling under competitive pressure from open-weight rivals, while the underlying compute and memory that inference actually runs on is getting more expensive, not less, because AI demand itself is straining the physical supply chain. That's a genuinely different dynamic than the software industry's usual cost curve, where falling prices to customers have historically tracked falling underlying infrastructure costs.

For investors in AI infrastructure and model companies, the practical question is which labs can absorb margin compression from both directions long enough to hold pricing power, and which get squeezed into raising prices again or slowing the pace of future cuts. Labs with their own chip supply commitments or long-term memory contracts locked in ahead of this squeeze are in a materially better position than those buying capacity on the spot market. What to watch: whether OpenAI, Google or DeepSeek adjust pricing again as memory and packaging costs continue climbing through the rest of the year.

ShareXLinkedInEmail
More onOpenAI →

Analysis and editorial commentary by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.

Read Next

AI· Aug 1, 2026

OpenAI's Astra Cracks 10 Unsolved Math Problems

Illustration for: OpenAI's Astra Cracks 10 Unsolved Math Problems
AI

OpenAI's Astra Cracks 10 Unsolved Math Problems

An internal version of OpenAI's next model, Astra, solved ten previously-open math and computer-science problems for roughly $2,000 in compute -- proofs Fields Medalist Timothy Gowers says he'd back for a top journal.

AI· Aug 3, 2026

AI's Dual-Use Risk Just Became a Live Security Problem

Illustration for: AI's Dual-Use Risk Just Became a Live Security Problem
AI

AI's Dual-Use Risk Just Became a Live Security Problem

A scoping error that let Claude models breach real production systems, and a US-China robot-ban standoff threatening rare-earth retaliation, broke days apart -- proof AI's dual-use risk is now operational, not hypothetical.

AI· Aug 3, 2026

DeepSeek's Offensive Use Forces a Red-Team Rethink

Illustration for: DeepSeek's Offensive Use Forces a Red-Team Rethink
AI

DeepSeek's Offensive Use Forces a Red-Team Rethink

Palo Alto Networks caught an operator using DeepSeek to autonomously attack 460+ systems after Claude and OpenAI's models refused the same job -- a live test of whether model-level refusal is a real safety layer or just a routing problem.

@Trace_Cohen·t@nyvp.com