Gimlet Labs is now worth $3 billion β more than triple its price six months ago β after a $300 million Series B led by Andreessen Horowitz.
Gimlet Labs sells a "multi-silicon inference cloud" that runs AI models across several kinds of chips instead of just one. On September 4, 2026, the three-year-old, San Francisco-based company announced a $300 million Series B that values it at $3 billion β up from roughly $980 million just six months earlier. The round arrives as the company says it has locked in billions of dollars of contracted revenue from customers that include one of the world's three largest frontier AI labs and one of the three largest cloud hyperscalers, though neither the labs nor the exact contract terms have been named publicly.

Gimlet Labs Valuation 2026: The $300 Million Series B and What $3 Billion Buys
Andreessen Horowitz led Gimlet Labs' $300 million Series B, which closed September 4, 2026 at a $3 billion post-money valuation, according to the company's own funding announcement. Existing investors Sapphire Ventures, Menlo Ventures, and Factory all returned, joined by two new backers with strategic weight: M12, Microsoft's corporate venture arm, and Arm, the chip design firm whose instruction-set architecture underpins much of the mobile and increasingly server chip market.
The company has not disclosed an exact Series A valuation, but Tech Funding News reported the $80 million Series A priced the company at roughly $980 million in March 2026, meaning the Series B represents better than a 3x markup in about six months β a compressed timeline even by 2026 AI-funding standards.
From Stanford Research Project to $3 Billion in Three Years
Gimlet Labs was founded in 2023 by Zain Asgar, who serves as CEO, alongside Michelle Nguyen, Omid Azizi, Natalie Serrino, and James Bartlett. The founding team previously worked together at Pixie Labs, and the company's core technology grew out of a Stanford University research project on splitting AI compute workloads across heterogeneous hardware. Gimlet stayed out of public view until it emerged from stealth in October 2025, when it disclosed its Series A alongside eight-figure annualized revenue.
Headcount and infrastructure scale moved just as fast: the company says it tripled its customer base between its October 2025 stealth exit and its March 2026 Series A, adding one of the top three frontier AI labs and one of the top three cloud hyperscalers as customers in that window, per its Series A announcement. By the Series B, Gimlet said it is scaling its managed infrastructure footprint toward several hundred megawatts of heterogeneous compute capacity.
One read on the founding thesis: before Gimlet existed, frontier labs and hyperscalers with enough engineering headcount were already building similar routing tooling in-house β splitting inference workloads across whatever chips they had on hand rather than buying a packaged product, since no vendor sold that layer commercially. Gimlet's bet is that turning that in-house tooling into a standalone managed service, sold to labs and enterprises that would rather buy than build it themselves, is a large enough market to support an independent company rather than a feature every hyperscaler eventually ships for free.
What "Multi-Silicon Inference" Actually Means
Running an AI model in production splits into two rough phases: "prefill," where the model reads and processes a prompt, and "decode," where it generates output tokens one at a time. The two phases have very different hardware needs β prefill is compute-heavy and benefits from raw throughput, while decode is memory-bandwidth-heavy and benefits from fast, cheap memory access. Most inference platforms run both phases on the same GPU fleet, typically Nvidia's. Gimlet's software disaggregates the two phases and routes each to whichever chip architecture handles it most cost-effectively β potentially Nvidia GPUs for one phase and custom silicon such as Google TPUs, AMD accelerators, or Amazon's Trainium/Inferentia chips for the other, according to reporting on the company's six-chip strategy.
Gimlet describes the resulting product, Gimlet Cloud, as the industry's first multi-silicon inference cloud built specifically for agentic AI workloads β chains of many small model calls, tool uses, and retries that behave differently from a single chatbot response, according to the company's own Series B announcement. The pitch to customers is straightforward: because agentic workloads chain together many inference calls rather than one, small per-call efficiency gains compound quickly across an entire agent session, and the company says this approach improves latency, throughput, hardware utilization, and power efficiency versus single-architecture serving.
The Inference Market Gimlet Is Selling Into
Gimlet's timing lines up with a real inflection in how enterprises spend on AI. Gartner forecasts worldwide AI-optimized infrastructure-as-a-service spending will grow 96% in 2026 to $42.3 billion, and β for the first time β global spending on inference ($23.3 billion) will exceed spending on training ($19 billion) this year. That shift, from paying to build models to paying to run them at scale, is the exact budget line Gimlet, Baseten, Fireworks AI, and Together AI are all competing for.
The practical problem Gimlet is pointing at is real: Nvidia GPUs remain the default choice for both prefill and decode phases of inference even though the two phases have different hardware requirements, largely because building a routing layer across chip architectures from Nvidia, AMD, Google, and Amazon is hard engineering that most AI labs and enterprises would rather buy than build in-house. Whether that routing layer is worth a standalone $3 billion company, versus a feature Nvidia, AMD, or the hyperscalers themselves eventually ship natively, is the open question the next funding round has to answer.
Gimlet Labs vs. the Rest of the AI Inference Market
Gimlet's $3 billion price sits well below several inference-serving peers that have disclosed harder revenue numbers over the past year. Fireworks AI and Baseten both crossed multibillion-dollar valuations on the back of eight- and nine-figure ARR figures they made public; Gimlet has not published a comparable number, framing its scale instead around "billions of dollars in contracted revenue" β signed multi-year commitments rather than revenue already collected and recognized.
| Company | Valuation | Disclosed revenue | Positioning |
|---|---|---|---|
| Fireworks AI | $17.5B | ~$800M annualized (May 2026) | Custom model serving, deep customization |
| Baseten | ~$13B | ~$600M annualized (Q1 2026) | Managed inference, serving reliability |
| Together AI | $8.3B | Not fully disclosed | Open-source model serving, price competition |
| Gimlet Labs | $3B | "Billions" contracted, not ARR | Multi-silicon, disaggregated agentic inference |
| Groq | $3.5B | Not fully disclosed | Custom LPU chips, deterministic low latency |
| Modal | ~$1.1B (2025) | Not fully disclosed | Serverless GPU compute for developers |
Figures from Sacra, Tech Funding News, TechCrunch, company disclosures, and prior Value Add VC reporting, as of September 2026. Together AI and Groq revenue not fully disclosed as of publication.
The comparison exposes the real gap in Gimlet's story: every peer above it on valuation has published a specific ARR or annualized-revenue figure, while Gimlet's largest disclosed number is a contracted-revenue total with no stated timeframe or recognition schedule. a16z is underwriting the multi-silicon architecture bet and the two blue-chip customer logos (undisclosed by name) more than a proven revenue base at this stage.
Why a16z Is Betting on Chip Diversity, Not Just More GPUs
Andreessen Horowitz's thesis, echoed across the deal's press coverage, is that the inference market is shifting from "buy more Nvidia GPUs" toward "route each workload to its most efficient chip," as hyperscalers and frontier labs diversify away from single-vendor dependence with custom silicon like Google's TPUs, Amazon's Trainium, and AMD's MI-series accelerators. Arm's participation as a new investor is notable given the firm's own push into server and AI chip licensing β a strategic, not purely financial, stake in whichever software layer ends up routing workloads across the widest range of chip architectures.
This likely means Gimlet is being priced less on current revenue and more on the credibility of its founding team's Stanford-derived disaggregation research and its two marquee customer relationships. That is a materially different bet than the one underwriting Fireworks AI's or Baseten's valuations, where published ARR growth carries more of the argument.
What the Headline Misses
"Billions of dollars in contracted revenue" is a backlog figure, not recognized revenue β the same distinction that has drawn scrutiny toward other AI infrastructure players who lead with multi-year contract totals instead of an ARR number. Gimlet has not said how many years those contracts span, when the revenue converts to cash, or what its actual trailing annualized revenue is today, which makes a direct comparison to Fireworks AI's disclosed ~$800 million annualized figure or Baseten's ~$600 million impossible on the numbers Gimlet has made public.
There's also a concentration risk built into the story: Gimlet's headline growth narrative rests on landing exactly two large, unnamed customers β one frontier lab, one hyperscaler β rather than a broad base of paying accounts. Fireworks AI, by contrast, has named more than 10,000 customers including Cursor, Perplexity, Notion, and DoorDash. If either of Gimlet's two anchor customers builds the disaggregation capability in-house or switches providers, the growth story underpinning a $3 billion valuation would need a different set of proof points.
The Bull and Bear Case
Bull case: Gimlet has landed two of the hardest customer logos in AI β a top-three frontier lab and a top-three hyperscaler β without a large existing sales organization, and its technical bet (that agentic workloads reward chip-level disaggregation) addresses a real, growing inefficiency as inference spend scales faster than any single chip vendor's supply. Arm and Microsoft's M12 joining as strategic investors is a signal that at least two large infrastructure players see enough architectural upside to take a direct stake rather than simply becoming customers.
Bear case: a $3 billion valuation on undisclosed audited revenue, a two-customer concentration story, and a "contracted revenue" figure with no stated recognition timeline asks investors to trust the pitch more than the numbers. Every inference-serving peer priced higher than Gimlet β Fireworks AI, Baseten, Together AI β has published a specific ARR figure; Gimlet has not, which is either a sign the number isn't yet impressive enough to lead with, or simply a different disclosure preference. Either way, it's the gap a $3 billion price has to close before the next round.
The Bottom Line
Gimlet Labs' $3 billion valuation is a bet on an architecture, not yet a disclosed revenue base: that routing AI inference across multiple chip types, rather than optimizing serving on one, becomes the default as agentic workloads scale and hyperscalers diversify away from single-vendor GPU dependence. The two customer logos and the Arm and M12 investments back that bet with real strategic weight. But every peer priced above Gimlet in this market has already shown its ARR math in public, and Gimlet's next round will likely need to do the same.
Track valuation multiples across the AI sector on the AI Valuations dashboard at Value Add VC. Reach out at t@nyvp.com or @Trace_Cohen.
Latest from the Pulse
Get VC data most people never see
β 100% free
Weekly benchmarks, valuations, and fund data. Join 5,000+ investors. No spam.