Analysis
McKinsey's State of AI in 2026 survey found that while per-token AI model costs have collapsed, enterprise AI spending has risen even faster, according to Fortune.
The Core Paradox
Token costs for models with roughly GPT-4-class benchmark performance fell from about $20 per million output tokens to as low as 7 cents through 2024, citing Stanford HAI data -- a decline of more than 280 times. Yet enterprise LLM spend tripled over the following 12 months. The explanation is consumption, not price: 40% of large organizations, those with more than $1 billion in annual revenue, now report scaling AI agents in production, up from 27% a year earlier, and agentic workflows consume far more tokens per completed task than a single chat response ever did -- multi-step reasoning, retries, tool calls and retrieval all add up.
“Yet enterprise LLM spend tripled over the following 12 months.”
The CFO Reaction
93% of survey respondents reported blowing through their token budgets, and McKinsey's own CFO is quoted describing the expense as "becoming a really big expense... going up at a vertical level." That reaction marks a shift from earlier 2020s-era AI cost conversations, which mostly focused on whether AI was affordable at all, to a more mature 2026 conversation about whether AI spend is being managed with the same rigor as any other fast-growing line item.
What This Means For Vendors And Buyers
The finding directly complicates how AI infrastructure and application vendors pitch cost savings: falling per-token prices, the metric most model providers lead with in their own marketing, does not translate into falling total bills once agentic usage scales, and enterprise buyers are increasingly sophisticated enough to know the difference. For startups selling into enterprises on a pure cost-savings pitch, the more durable argument is now about managing and predicting consumption -- not just accessing cheaper tokens -- which is exactly the problem this week's other funding stories, from Baselayer's agent-identity infrastructure to CoreWeave's compute financing, are each addressing from different angles of the same underlying spend-management challenge.