Analysis
Snorkel AI, the Redwood City-based AI data company, raised a $350 million Series E co-led by Insight Venture Management and Section 32, valuing the company at $3.5 billion, according to PR Newswire and TechCrunch. Returning investors Addition, Lightspeed, Greylock, GV, P7, Wells Fargo, Walden Catalyst Ventures and Factory all participated, alongside new backers Third Point, March Capital, Blumberg Capital, Allegis, Standard VC and Frontline.
The new valuation is nearly triple where Snorkel stood 17 months ago, when it closed a much smaller Series D.
From Weak Supervision To A Data Factory
Snorkel was founded in 2019 as a Stanford spinout built around the open-source Snorkel project for programmatically labeling training data. About a year ago the company pivoted from selling software licenses to running what it now calls an "agentic data factory" -- pairing AI tooling with a managed expert network to build the complex training data, reinforcement-learning environments and evaluation rubrics that even skilled human specialists take hours or days to construct. That data-as-a-service offering is what grew 18x and pushed Snorkel's annualized revenue run rate past $375 million in the week of this announcement, a seven-year-old company disclosing figures unusually specific for the category.
The Competitive Field
Snorkel competes for frontier-lab training-data budget against Scale AI, Surge AI, Turing and Mercor, each pursuing a different mix of automation and managed human expert networks. Snorkel's $3.5 billion mark sits well below Scale AI's valuation, but the disclosed $375 million ARR figure gives it a revenue-to-valuation ratio -- close to 1x run-rate revenue on this raise -- that is unusually grounded for an AI infrastructure company at this stage, where most peers still raise primarily against committed capacity or customer pipeline rather than booked revenue.
Snorkel's growth is itself a bet that frontier labs' own price war over inference won't extend to what they spend on training data -- if anything, cheaper inference funds more experimentation, and more experimentation needs more high-quality data. The risk investors are underwriting at this multiple is customer concentration: frontier labs are also the companies most likely to eventually build equivalent data-labeling operations in-house, as OpenAI, Anthropic and Meta have all begun doing to varying degrees.
The open question is how much of that $375 million run rate sits inside three or four frontier-lab contracts that could each choose to build the same capability internally once they have enough scale and in-house expertise to justify it.