$600 million is Baseten's annualized revenue run-rate as of March 2026 โ up from $200 million just three months earlier in December 2025, a nearly 1,900% year-over-year jump. That's the short answer. The longer answer is that Baseten never charges customers a flat license fee at all โ it bills by the GPU-minute for dedicated model deployments and per-token for its open-source Model APIs, and that usage-based mechanic is exactly what took its valuation from $2.15 billion to $13 billion in under nine months.
Baseten sells production infrastructure for AI inference โ the unglamorous, expensive layer that sits between a trained model and a live product feature. Every prompt a customer's app sends to a fine-tuned or open-source model gets routed through Baseten's orchestration layer, which is why revenue scales almost mechanically with how much AI product usage its customers see. Here's exactly how the money works.
Figures compiled from TechCrunch, Sacra, Dealroom, and Digital Applied reporting on Baseten's revenue and funding, as of July 2026.
How does Baseten make money
Baseten makes money by charging enterprises to run AI inference โ serving a trained model's predictions in production โ rather than charging per employee seat like traditional SaaS. Dedicated model deployments are billed by the GPU-minute, with published 2026 rates spanning $0.01052 per minute on a basic T4 GPU up to $0.02012 on an A10G, $0.06667 on an A100 80GB, and $0.16633 on a top-tier B200, with idle replicas scaling to zero and costing nothing.
The second revenue line is Model APIs, a per-token catalog for popular open-source models that lets customers skip infrastructure setup entirely and pay only for tokens generated, similar to how OpenAI or Anthropic bill their own APIs. Enterprise accounts layer on custom SLAs, private or hybrid deployment options, and negotiated compute discounts, but the core mechanic stays the same across every tier: the bill is a direct function of compute consumed, not seats provisioned, which is why revenue compounds automatically as a customer's product usage grows without an additional sales conversation.
Baseten's valuation: from $2.15B to $13B in nine months
Baseten closed a $150 million Series D in September 2025 at a $2.15 billion valuation led by BOND. Four months later, in January 2026, it raised a $300 million Series E at a $5 billion valuation led by IVP and CapitalG, with Nvidia joining as a strategic investor โ a signal that the chipmaker wanted exposure to the inference layer, not just the training layer, of the AI stack.
By June 2026, Baseten closed a $1.5 billion Series F at a $13 billion valuation led by Altimeter Capital, Conviction Partners, and Spark Capital โ roughly 6x the January valuation in five months. At $600 million trailing ARR, that prices Baseten at roughly 22x revenue, a steep but not unusual multiple in a market where AI infrastructure companies with triple-digit growth routinely trade above 20x. For more on how growth-stage AI companies get priced, see our AI valuations dashboard.
| Round | Date | Amount Raised | Valuation | Lead Investors | Implied ARR Multiple |
|---|---|---|---|---|---|
| Series D | Sep 2025 | $150M | $2.15B | BOND | ~36x (est. $60M ARR) |
| Series E | Jan 2026 | $300M | $5B | IVP, CapitalG, Nvidia | ~14x ($350M ARR) |
| Series F | Jun 2026 | $1.5B | $13B | Altimeter, Conviction, Spark | ~22x ($600M ARR) |
| Together AI Series C | Jul 2026 | $800M | $8.3B | Undisclosed | ~7x ($1.15B bookings) |
| Fireworks AI (last known) | 2025 | Undisclosed | Not disclosed | Undisclosed | N/A |
| Modal (last known) | 2025 | Undisclosed | Not disclosed | Undisclosed | N/A |
Figures blended from TechCrunch, Pulse2, Dealroom, and Sacra reporting on Baseten and Together AI funding rounds. ARR multiples are estimates based on the nearest disclosed revenue figure at each round.
Baseten vs Modal, Together AI, and Fireworks AI
Baseten competes in the AI inference infrastructure category against Together AI, Fireworks AI, and Modal โ all four companies rent GPU capacity across cloud providers and resell it to enterprises that need to serve trained models reliably at scale, rather than owning their own data centers. As of mid-2026, Together AI leads on scale with $1.15 billion in annual bookings after an $800 million Series C at an $8.3 billion valuation, followed by Fireworks AI at roughly $800 million in annualized revenue, Baseten at $600 million, and Modal at $300 million ARR.
The differentiation between them comes down to developer workflow rather than raw pricing, since all four use similar usage-based, per-minute or per-token billing. Baseten's edge is its open-source Truss framework for packaging custom models into managed endpoints, which is why it's disproportionately popular with ML teams deploying fine-tuned or proprietary models rather than just calling a hosted open-source API. See our SaaS valuations dashboard for how usage-based infrastructure companies get priced relative to seat-based software.
Why usage-based pricing scales faster than seat-based SaaS
Baseten's revenue mix is roughly 60% dedicated GPU-minute deployments, 30% Model APIs per-token billing, and 10% enterprise support and custom SLAs โ a structure that means the company's top-line growth is directly tied to how much AI product usage its customers see, not how many logins they provision. That's the core reason inference infrastructure companies have grown faster in 2026 than almost any other software category: every additional AI feature a customer ships generates more inference calls automatically, with zero incremental sales motion required.
The tradeoff is that usage-based revenue is also more volatile than subscription SaaS โ if a customer's product usage drops, so does Baseten's bill from that account, with no minimum commitment floor unless it's negotiated into an enterprise contract. That volatility is part of why investors price these companies on revenue multiples closer to 15-25x rather than the 8-12x more typical of steady-state seat-based software, since the growth curve can bend sharply in either direction.
What Baseten's growth signals for AI infrastructure investing
Baseten's path from $60 million in estimated ARR at its Series D to $600 million eight months later is one of the fastest revenue scaling curves in enterprise software history, and it's happening in a category โ AI inference โ that barely existed as a standalone line item three years ago. That growth rate is also why four separate companies in the same narrow category (Baseten, Together AI, Fireworks AI, Modal) have all raised at multi-billion-dollar valuations within the same twelve-month window, a concentration of capital that mirrors what happened in cloud infrastructure a decade earlier.
For founders and LPs watching this space, the signal isn't just that inference is a big market โ it's that usage-based, infrastructure-layer businesses tied directly to AI product consumption are compounding revenue faster than almost any subscription software category has in the last cycle. That's a pattern worth tracking closely on our big tech earnings dashboard, since hyperscaler capex guidance is the leading indicator for how much inference demand companies like Baseten can capture next.
Who's actually buying Baseten's inference infrastructure
Baseten's customer base skews toward ML teams shipping proprietary or fine-tuned models rather than companies that just want a hosted API for a stock open-source model โ that's the segment where its open-source Truss packaging framework matters most, since it lets a team go from a trained model checkpoint to a production HTTPS endpoint without building custom serving infrastructure in-house. Enterprise accounts negotiate compute discounts and custom SLAs once their monthly spend crosses into six figures, and Baseten's asset-light approach โ renting capacity across more than 15 cloud providers instead of owning data centers โ means it can offer whichever GPU generation (T4 through B200) fits a given workload's latency and cost profile without waiting on its own hardware procurement cycle.
That multi-cloud GPU sourcing model is also Baseten's main defense against the two biggest risks facing every inference reseller right now: GPU scarcity pricing spikes and hyperscalers building competing first-party inference products. By not being locked into a single chip supplier or a single cloud's capacity, Baseten can shift workloads toward whichever provider has available inventory at the best price in a given week โ a flexibility that vertically integrated hyperscaler inference offerings, tied to their own data centers, structurally can't match at the same speed.
Bottom line: Baseten makes money by charging enterprises per GPU-minute for dedicated model deployments and per-token for its Model APIs catalog, a usage-based model that took ARR from $200 million to $600 million in a single quarter and pushed its valuation from $2.15 billion to $13 billion in under nine months. It now sits third among the four major AI inference infrastructure companies by revenue, behind Together AI's $1.15 billion in bookings and Fireworks AI's $800 million ARR, but ahead of Modal's $300 million โ a scoreboard that will likely keep shifting fast as hyperscaler AI capex continues to flow directly into inference demand.
Get VC data most people never see โ free.
Weekly benchmarks, valuations, and fund data. No spam, unsubscribe anytime.