VC
Value Add VC
⚡HomePulse⚡Helpful Apps📝Blog🤝Partner
Illustration for: Vals AI Raises $40M to Grade AI Models
Value Add VC/Pulse/AIDEEP DIVE$40M at $400M

Vals AI Raises $40M to Grade AI Models

San Francisco-based Vals AI raised a $40 million Series A led by a16z at a $400 million valuation, building independent benchmarks that test frontier AI models on real professional tasks rather than academic leaderboards.

By the Numbers

$40M
New round
$400M
New valuation
8x
Revenue growth vs. 2025
~52%
Frontier model fail rate (finance tasks)
TC
By the AI Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
August 13, 2026
2 min read
ShareXLinkedInEmail

THE RUNDOWN

1

Vals AI raised $40 million in a Series A round led by a16z at a $400 million valuation, with participation from existing investors 8VC and Bloomberg Beta plus new backers HRT Ventures and Next Ladder Ventures, per [Tech Funding News](https://techfundingnews.com/a16z-leads-40m-vals-ai-round-at-400m-valuation-to-test-ai-on-real-world-tasks/)

2

The company builds independent evaluation benchmarks that test frontier models against real-world professional tasks, and its research has found frontier models fail roughly 52% of real finance-analyst tasks it evaluates

3

Revenue has grown eightfold compared with all of 2025, the customer base has doubled and the team has tripled over the past six months, as enterprises, labs and governments seek independent measurement of model ROI

4

The round underscores how model evaluation has become its own funded category as buyers grow skeptical of vendor-reported benchmarks -- the same skepticism that has dogged self-reported figures from OpenAI, Anthropic and Google throughout this year's model release cycle

TC

The VC Read · Trace's Take

Trace Cohen

The 52% finance-task failure rate is the number every enterprise buyer evaluating an AI vendor pitch should be asking for before signing, not after -- if I'm diligencing a fintech or legal-AI startup right now, I want their Vals AI score in the data room, not their marketing benchmark. Independent evaluation becoming a funded category in its own right is a healthy sign for the ecosystem; the risk is that labs eventually route around independent evaluators the way they've resisted independent audits of training data.

AI Landscape →

Analysis

Vals AI, a San Francisco startup that builds independent benchmarks for testing AI models on real-world professional tasks, raised $40 million in a Series A round led by a16z at a $400 million valuation, according to Tech Funding News. Existing investors 8VC and Bloomberg Beta returned, joined by new investors HRT Ventures and Next Ladder Ventures.

Vals AI's pitch runs counter to how most AI labs market model capability -- rather than academic leaderboards or lab-published benchmark scores, the company evaluates models against tasks pulled from actual professional workflows: finance, law, and other domains where errors carry real financial or compliance consequences. Its research has found that frontier models fail roughly 52% of the real finance-analyst tasks it tests them against, a figure that stands in sharp contrast to the near-saturated scores most frontier labs report on standard academic benchmarks like MMLU or GPQA.

Independent measurement as its own category

The company's growth numbers are the strongest signal in the round: revenue has grown eightfold compared with all of 2025, its customer base has doubled, and its team has tripled over the past six months. That trajectory reflects a structural shift in how enterprises, labs and governments are buying AI -- self-reported benchmark scores from OpenAI, Anthropic, Google and the rest of the frontier field have become less trusted as a basis for procurement decisions, particularly after several labs faced criticism this year for benchmark methodology that inflated real-world capability claims.

Vals AI's bet is that independent, task-specific evaluation becomes required infrastructure for any enterprise deploying AI in a regulated or high-stakes function, the same way independent auditors became required infrastructure for financial reporting once self-reported numbers alone stopped being sufficient for institutional trust. Competing directly in this category are a handful of smaller evaluation startups and academic benchmark consortia, none of which have yet raised at a comparable valuation -- Vals AI's round effectively establishes it as the best-funded independent player in AI evaluation heading into a year where enterprise AI spending scrutiny is intensifying rather than easing.

ShareXLinkedInEmail

More on

Andreessen Horowitz →

Reported by Tech Funding News · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.

Read Next

AI· Aug 17, 2026

Cursor Launches Origin to Take On GitHub

Illustration for: Cursor Launches Origin to Take On GitHub
AI

Cursor Launches Origin to Take On GitHub

Cursor's maker Anysphere launched Origin, an AI-native code hosting platform built into the editor, the same week a major GitHub outage exposed how much of the AI coding stack leans on a single hosting layer.

AI· Aug 17, 2026

Anthropic's Annualized Revenue Hits $65B in July

Illustration for: Anthropic's Annualized Revenue Hits $65B in July
AI$65B annualized run rate

Anthropic's Annualized Revenue Hits $65B in July

Anthropic told investors its annualized revenue run rate climbed to $65 billion at the end of July, a sevenfold jump from about $9 billion at the end of 2025, as it prepares for an IPO expected this fall.

AI· Aug 18, 2026

MIT Finds AI Models Develop 'Amnesia' at Scale

Illustration for: MIT Finds AI Models Develop 'Amnesia' at Scale
AI

MIT Finds AI Models Develop 'Amnesia' at Scale

MIT researchers found that as generative AI models grow larger, their outputs become nearly impossible to trace back to specific training examples -- a phenomenon they call attribution decay that complicates copyright and fair-use fights over AI-generated.

@Trace_Cohen·t@nyvp.com