Illustration for: McKinsey: AI Got 280x Cheaper, Enterprise Bills Tripled Anyway

McKinsey: AI Got 280x Cheaper, Enterprise Bills Tripled Anyway

McKinsey found per-token AI model costs fell from roughly $20 to as low as 7 cents per million tokens through 2024, yet enterprise LLM spending tripled the following year as 93% of surveyed companies blew through their token budgets.

By the Numbers

>280x
Token cost decline
3x
Enterprise LLM spend, YoY
93% of firms
Over token budget
40%, up from 27%
Large orgs scaling agents
TC
By the AI Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
2 min read
ShareXLinkedInEmail

THE RUNDOWN

1

Token costs for roughly GPT-4-class model performance fell from about $20 per million output tokens to as low as 7 cents -- a decline of more than 280x -- citing Stanford HAI data, yet enterprise LLM spend tripled over the following 12 months rather than falling.

2

93% of surveyed respondents reported exceeding their token budgets, and CFOs increasingly describe AI as a rapidly escalating expense line rather than a cost center trending toward zero, directly contradicting the assumption that cheaper tokens mean cheaper AI adoption overall.

3

40% of large organizations (over $1 billion in annual revenue) now report scaling AI agents in production, up from 27% a year earlier -- agent adoption, which multiplies token consumption per task, is the primary driver of rising spend even as per-token prices fall.

4

The paradox directly complicates how AI-native startups should pitch cost savings to enterprise buyers: cheaper unit economics does not automatically mean a lower total bill once usage scales with agentic workflows, and CFOs are increasingly aware of that gap.

TC

The VC Read · Trace's Take

Trace Cohen

Every AI-native pitch deck built around 'our costs will fall as models get cheaper' has this paradox baked into it and most founders haven't priced that in -- consumption scales faster than unit costs fall the moment agentic workflows get involved, which means the real product to build might be spend management, not cheaper inference. This is the single most useful data point in this issue for anyone underwriting AI-application gross margins over a two-year hold.

Analysis

McKinsey's State of AI in 2026 survey found that while per-token AI model costs have collapsed, enterprise AI spending has risen even faster, according to Fortune.

The Core Paradox

Token costs for models with roughly GPT-4-class benchmark performance fell from about $20 per million output tokens to as low as 7 cents through 2024, citing Stanford HAI data -- a decline of more than 280 times. Yet enterprise LLM spend tripled over the following 12 months. The explanation is consumption, not price: 40% of large organizations, those with more than $1 billion in annual revenue, now report scaling AI agents in production, up from 27% a year earlier, and agentic workflows consume far more tokens per completed task than a single chat response ever did -- multi-step reasoning, retries, tool calls and retrieval all add up.

Yet enterprise LLM spend tripled over the following 12 months.

The CFO Reaction

93% of survey respondents reported blowing through their token budgets, and McKinsey's own CFO is quoted describing the expense as "becoming a really big expense... going up at a vertical level." That reaction marks a shift from earlier 2020s-era AI cost conversations, which mostly focused on whether AI was affordable at all, to a more mature 2026 conversation about whether AI spend is being managed with the same rigor as any other fast-growing line item.

What This Means For Vendors And Buyers

The finding directly complicates how AI infrastructure and application vendors pitch cost savings: falling per-token prices, the metric most model providers lead with in their own marketing, does not translate into falling total bills once agentic usage scales, and enterprise buyers are increasingly sophisticated enough to know the difference. For startups selling into enterprises on a pure cost-savings pitch, the more durable argument is now about managing and predicting consumption -- not just accessing cheaper tokens -- which is exactly the problem this week's other funding stories, from Baselayer's agent-identity infrastructure to CoreWeave's compute financing, are each addressing from different angles of the same underlying spend-management challenge.

ShareXLinkedInEmail

Key Sources

2 sources

Reported by Fortune · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.