VC
Value Add VC
⚡HomePulse⚡Helpful Apps📝Blog🤝Partner
Illustration for: Why Multi-Model Routing Just Became the Default Architecture
Value Add VC/Pulse/AI

Why Multi-Model Routing Just Became the Default Architecture

From Fireworks' inference business to Microsoft's Project Perception security tool, this week's biggest AI stories share one underlying architecture: routing tasks across multiple models by cost and fit rather than defaulting to a single frontier model.

By the Numbers

95%+ of tokens
Fireworks fine-tuned share
3 providers routed
Project Perception model count
$15/1M tokens
Kimi K3 output cost
TC
By the AI Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
July 19, 2026
1 min read
ShareXLinkedInEmail

THE RUNDOWN

1

Fireworks says more than 95% of the 40 trillion tokens it serves daily come from models fine-tuned and routed for specific customer workloads, not general-purpose frontier-model calls -- the core of its $17.5 billion valuation

2

Microsoft's Project Perception explicitly routes security-scanning tasks across Microsoft, OpenAI and Anthropic models to reserve expensive frontier calls for only the steps that need them

3

Kimi K3's aggressive pricing -- $15 per million output tokens for near-frontier performance -- gives infrastructure companies another cheap option to route lower-stakes tasks toward, accelerating the shift away from single-model dependence

4

The pattern suggests the next competitive battleground in AI infrastructure isn't which lab has the best frontier model, but which platform routes most efficiently across all of them

TC

The VC Read · Trace's Take

Trace Cohen

Every founder still architecting a product around a single model provider is building on a foundation that the market just told you is obsolete -- Fireworks, Microsoft and Kimi K3's pricing all point the same direction. The moat isn't picking the best model anymore, it's building the routing layer that picks the best model for you, task by task, as pricing and capability shift weekly. If your AI product doesn't have a model-abstraction layer by now, you're one aggressive competitor's pricing move away from a very expensive rewrite.

Analysis

Three of this week's biggest AI stories -- Fireworks' $17.5 billion valuation, Microsoft's Project Perception, and Kimi K3's aggressive pricing -- point to the same underlying shift: enterprises are moving away from defaulting every task to a single frontier model and toward routing tasks across multiple models by cost and fit. Fireworks' entire business is built on this: more than 95% of the 40 trillion tokens it serves daily come from models fine-tuned for specific customer workloads rather than general-purpose API calls to a foundation model. Microsoft's Project Perception applies the same logic to security scanning, assigning tasks to Microsoft's, OpenAI's or Anthropic's models depending on which is cheapest and best-suited for that specific step.

The economics driving this are straightforward. Frontier-model API calls from OpenAI, Anthropic or Google carry a meaningful premium over smaller, specialized, or lower-cost alternatives -- and Kimi K3's arrival this week, at $15 per million output tokens for near-frontier benchmark performance, only widens the gap enterprises can exploit by routing intelligently rather than paying frontier prices for every task regardless of difficulty.

“This has real implications for how the AI infrastructure market gets valued.”

This has real implications for how the AI infrastructure market gets valued. If routing architecture becomes the default, the companies that win aren't necessarily the labs with the single best model -- they're the platforms that can reliably route a given task to whichever model handles it best at the lowest cost, a category currently being contested by Fireworks, Together AI, SambaNova, and now implicitly by Microsoft's own enterprise stack.

For founders building AI products, the practical takeaway is that hard-coding a single model provider into your architecture is an increasingly risky decision -- both because pricing keeps shifting the economics and because new entrants like Kimi K3 keep changing which model is actually the best choice for a given task. What to watch next: whether routing-layer companies start commanding valuation premiums over pure model labs as this architecture becomes standard practice.

ShareXLinkedInEmail

More on

Microsoft →

Reported by Value Add Pulse Analysis · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.

Read Next

AI· Aug 14, 2026

OpenAI Sheds Senior Execs in Pre-IPO Shakeup

Illustration for: OpenAI Sheds Senior Execs in Pre-IPO Shakeup
AI

OpenAI Sheds Senior Execs in Pre-IPO Shakeup

OpenAI has lost its chief revenue officer, its longtime COO and several senior leaders within days of each other, as co-founder Greg Brockman consolidates operating control ahead of a planned public listing.

AI· Aug 13, 2026

Anthropic's CFO Starts Courting IPO Investors

Illustration for: Anthropic's CFO Starts Courting IPO Investors
AI

Anthropic's CFO Starts Courting IPO Investors

Anthropic CFO Krishna Rao has begun early, informal meetings with prospective IPO investors, though he has not discussed valuation -- the $2 trillion figure circulating on Wall Street comes from investors' own math, not from Anthropic.

AI· Aug 13, 2026

Gemini 3.7 Flash Launches With 50% Price Cut for Coding

Illustration for: Gemini 3.7 Flash Launches With 50% Price Cut for Coding
AI

Gemini 3.7 Flash Launches With 50% Price Cut for Coding

Google released Gemini 3.7 Flash just three weeks after 3.6 Flash, cutting introductory API pricing in half while improving coding, debugging and enterprise-automation benchmarks over its predecessor.

@Trace_Cohen·t@nyvp.com