Analysis
Gimlet Labs, a 30-person startup that lets AI workloads shift seamlessly between Nvidia, AMD, and custom chips, raised $300 million in a Series B that values the company at $3 billion -- a roughly 33x jump from the $92 million in total funding it had raised as of March.
Andreessen Horowitz led the round, joined by existing backer Menlo Ventures and two new strategic investors: Arm Holdings and Microsoft's M12 venture arm, according to Bloomberg and the company's own announcement.
Founded in 2023 by CEO Zain Asgar -- a former Nvidia GPU architect and Google AI engineering lead -- alongside Michelle Nguyen, Omid Azizi, Natalie Serrino, and James Bartlett, Gimlet spent two years in stealth before launching publicly last October with what it calls the industry's first multi-silicon inference cloud. The pitch: instead of locking a workload to one vendor's GPUs, Gimlet's software splits an inference job across whatever mix of Nvidia, AMD, and custom accelerators is cheapest and fastest at that moment, claiming 3x-to-10x speedups at the same cost.
“Microsoft's M12, meanwhile, gives Gimlet a foothold inside Azure's own compute-allocation conversations.”
The chip-agnostic inference land grab
Gimlet isn't alone in betting that inference orchestration -- not raw compute -- is the next infrastructure layer worth owning.
Baseten closed a $1.5 billion Series F this summer on similar chip-neutral positioning, Fireworks AI raised $1.505 billion at a $17.5 billion valuation in July while serving 40 trillion tokens daily, and Together AI has built a comparable multi-cloud inference business.
What differentiates Gimlet is scale of ambition relative to size: at $3 billion, it's valued at roughly a fifth of Fireworks despite a fraction of the revenue and public token-serving numbers those rivals disclose.
The round also lands five months, not years, after Gimlet's $80 million Series A -- a pace that mirrors the broader 2026 AI infrastructure market, where multiple rounds inside a single calendar year are becoming the norm for vendors touching the GPU-scarcity problem directly:
- Volta -- raised a large round at a rich multi-billion-dollar valuation in August, backed by Nvidia and Michael Dell, to bankroll AI compute for smaller labs.
- Crusoe -- just closed a fresh multibillion-dollar raise months after its prior round, converting stranded-gas compute into an AI cloud business.
Money is moving faster than these companies can spend it -- a dynamic Gimlet's own back-to-back rounds fit squarely inside.
Arm's participation is the most strategically loaded detail. Arm-based server chips from Nvidia's Grace platform and Amazon's Graviton line are gaining share in AI inference specifically because they're cheaper per token than GPU-only stacks, and Arm has an obvious interest in software that makes its architecture a viable default rather than a niche option. Microsoft's M12, meanwhile, gives Gimlet a foothold inside Azure's own compute-allocation conversations.
That's also the risk sitting underneath the valuation: Gimlet's technology works by abstracting away chip differences, but the more successful it becomes, the more it invites Nvidia, AMD, or the cloud hyperscalers to build the same orchestration layer natively and cut out the middleman. Nvidia's own Dynamo inference framework and Google's Ironwood TPU stack already do pieces of what Gimlet sells as a standalone product -- a dependency risk no amount of chip-neutral positioning fully escapes. Gimlet says its roughly 30-person team and eight-figure revenue have tripled its customer base since October, but neither the company nor its investors have disclosed an actual ARR number alongside the new valuation.