Analysis
A new survey of enterprise AI spending finds the per-token list price and the actual realized cost of running AI at scale have become almost unrelated numbers. The LLM Token Expenditure Index fell to about $0.97 per million tokens this month, its lowest level since the index launched -- while enterprise AI bills at frontier-tier volume actually land between $3 and $12 per million tokens once real usage patterns are counted.
The gap is structural, not a pricing anomaly: agentic workflows -- retries, retrieval steps, orchestration overhead and observability logging -- multiply raw token consumption by 50 to 500 times relative to a single simple query, according to the survey. That means the advertised rate card functions as a floor on cost, not a reliable predictor of it. Two companies buying from the identical vendor at the identical list price can end up two to three times apart in cost per completed task, driven entirely by how they route requests, how large their prompts run, and how aggressively they retry failed calls.
“That means the advertised rate card functions as a floor on cost, not a reliable predictor of it.”
The measurement problem compounds the spending problem: only 20% to 25% of companies in McKinsey's May 2026 Enterprise AI FinOps survey reported having mature cost-management capabilities for their AI spend, and managing token cost and usage is now the single most commonly cited top challenge among FinOps teams. For any startup pitching AI-driven margin improvement, or any GP diligencing one, the realistic per-task cost figure -- not the list price the vendor quotes -- is the number that determines whether the unit economics actually work, and most companies still can't measure it with confidence.