VC
Value Add VC
โšกHomePulseโšกHelpful Apps๐Ÿ“Blog
โ† Value Add PulseAI

Why Multi-Model Routing Just Became the Default Architecture

From Fireworks' inference business to Microsoft's Project Perception security tool, this week's biggest AI stories share one underlying architecture: routing tasks across multiple models by cost and fit rather than defaulting to a single frontier model.

95%+ of tokens
Fireworks fine-tuned share
3 providers routed
Project Perception model count
$15/1M tokens
Kimi K3 output cost
TC
Trace Cohen
Early-stage VC & angel ยท Founder, New York Venture Partners
July 19, 2026
1 min read
ShareXLinkedInEmail
THE RUNDOWN
1

Fireworks says more than 95% of the 40 trillion tokens it serves daily come from models fine-tuned and routed for specific customer workloads, not general-purpose frontier-model calls -- the core of its $17.5 billion valuation

2

Microsoft's Project Perception explicitly routes security-scanning tasks across Microsoft, OpenAI and Anthropic models to reserve expensive frontier calls for only the steps that need them

3

Kimi K3's aggressive pricing -- $15 per million output tokens for near-frontier performance -- gives infrastructure companies another cheap option to route lower-stakes tasks toward, accelerating the shift away from single-model dependence

4

The pattern suggests the next competitive battleground in AI infrastructure isn't which lab has the best frontier model, but which platform routes most efficiently across all of them

TC
The VC Read ยท Trace's TakeTrace Cohen

Every founder still architecting a product around a single model provider is building on a foundation that the market just told you is obsolete -- Fireworks, Microsoft and Kimi K3's pricing all point the same direction. The moat isn't picking the best model anymore, it's building the routing layer that picks the best model for you, task by task, as pricing and capability shift weekly. If your AI product doesn't have a model-abstraction layer by now, you're one aggressive competitor's pricing move away from a very expensive rewrite.

Three of this week's biggest AI stories -- Fireworks' $17.5 billion valuation, Microsoft's Project Perception, and Kimi K3's aggressive pricing -- point to the same underlying shift: enterprises are moving away from defaulting every task to a single frontier model and toward routing tasks across multiple models by cost and fit. Fireworks' entire business is built on this: more than 95% of the 40 trillion tokens it serves daily come from models fine-tuned for specific customer workloads rather than general-purpose API calls to a foundation model. Microsoft's Project Perception applies the same logic to security scanning, assigning tasks to Microsoft's, OpenAI's or Anthropic's models depending on which is cheapest and best-suited for that specific step.

The economics driving this are straightforward. Frontier-model API calls from OpenAI, Anthropic or Google carry a meaningful premium over smaller, specialized, or lower-cost alternatives -- and Kimi K3's arrival this week, at $15 per million output tokens for near-frontier benchmark performance, only widens the gap enterprises can exploit by routing intelligently rather than paying frontier prices for every task regardless of difficulty.

โ€œThis has real implications for how the AI infrastructure market gets valued.โ€

This has real implications for how the AI infrastructure market gets valued. If routing architecture becomes the default, the companies that win aren't necessarily the labs with the single best model -- they're the platforms that can reliably route a given task to whichever model handles it best at the lowest cost, a category currently being contested by Fireworks, Together AI, SambaNova, and now implicitly by Microsoft's own enterprise stack.

For founders building AI products, the practical takeaway is that hard-coding a single model provider into your architecture is an increasingly risky decision -- both because pricing keeps shifting the economics and because new entrants like Kimi K3 keep changing which model is actually the best choice for a given task. What to watch next: whether routing-layer companies start commanding valuation premiums over pure model labs as this architecture becomes standard practice.

ShareXLinkedInEmail
More onMicrosoft โ†’

Originally reported by Value Add Pulse. Analysis and editorial commentary by Value Add Pulse.

โ† Back to Pulse

THE WIRE in your inboxโ€” Tech, startup & VC news with Trace's take. Free, no spam.

Read Next

AI$10B compute lease

Anthropic in Talks to Lease $10B of Meta's Compute

Anthropic and Meta are in early talks for Anthropic to lease up to $10B of Meta's AI compute over two years, letting Meta monetize its buildout while Anthropic diversifies beyond Amazon and Google.

AI

China's Open-Weight Wave Forces an Enterprise Rethink

Kimi K3's benchmark-topping debut is accelerating enterprise interest in open-weight models, forcing US buyers to weigh self-hosted Chinese models against closed, subscription-priced offerings from Anthropic and OpenAI.

AI$960,000

Jensen Huang's Leather Jacket Sells for $960K

A leather jacket worn by Nvidia CEO Jensen Huang sold for $960,000 at Sotheby's, nearly 20 times its pre-sale estimate, with proceeds benefiting a philanthropic initiative for young tech builders.

@Trace_Cohenยทt@nyvp.com