Illustration for: Alibaba's Own Benchmark: AI Agents Fail 40% of Tasks

Alibaba's Own Benchmark: AI Agents Fail 40% of Tasks

Alibaba's own commerce benchmark shows the strongest AI agent completing only 61.7% of real business tasks correctly, a gap its president says separates confident-sounding agents from agents that can actually execute work end to end.

TC
Trace Cohen
Early-stage VC & angel · Founder, New York Venture Partners
2 min read
ShareXLinkedInEmail
TC

The VC Read · Trace's Take

Trace Cohen

The real diligence question for any agent startup is not the demo -- it's the founder's own failure-rate benchmark on messy, real tasks; Alibaba just published one showing a 40% miss rate, and almost nobody else has.

Analysis

I spent part of this week reading Kuo Zhang's piece in Fortune, and it's a rare thing: an AI executive publishing his own company's benchmark showing his own company's agents failing. Zhang runs Alibaba.com, and his team built CommerceAgentBench, an open-source test of 107 real e-commerce tasks -- sourcing, landed-cost math, multi-leg shipping, after-sales disputes -- drawn from 10 million active small-business users, 1.6 million real conversations and 200,000 execution traces.

The headline number: the strongest frontier model Alibaba tested completed 61.7% of tasks. Put differently, nearly 40% of the time, a task that looked plausible in the transcript came back wrong when someone checked the actual output against the actual system. Zhang's framing is blunt -- agents can talk convincingly about landed costs, currency conversion and shipping routes, and still get the arithmetic wrong in a way that costs a real seller real money.

This matters beyond Alibaba because every major lab is currently selling some version of the opposite story:

The headline number: the strongest frontier model Alibaba tested completed 61.7% of tasks.

  • OpenAI -- ChatGPT's agent mode is marketed as able to book, buy and file on a user's behalf with minimal supervision.
  • Anthropic -- Claude's computer-use and "Cowork" features are pitched as production-ready for multi-step business workflows.
  • Google -- Gemini's Project Mariner demos browser agents completing shopping and research tasks autonomously.

None of those companies has published a CommerceAgentBench-style number showing their own agent's real-world task completion rate against a benchmark built from actual paying customers rather than curated demo tasks. Alibaba did, and it isn't flattering.

If I'm underwriting an AI agent startup right now, this changes what I ask for in diligence. A demo video or a cherry-picked eval score tells me almost nothing about whether the agent works on the tenth task, the messy one, the after-sales dispute with three PDFs attached. I want the founder's own failure-mode benchmark -- task categories, sample size, and the percentage that came back wrong -- not the percentage that came back impressive.

Room for disagreement: the 61.7% number is one company's internal benchmark on one company's task distribution, and Alibaba has an incentive to publish a number that makes room for a human-in-the-loop product tier it can upsell. Agent performance has also moved fast -- the frontier model that scored 61.7% today may not be the frontier model in three months, and coding agents already clear far higher completion rates on narrower, more structured tasks than open-ended commerce. It's possible CommerceAgentBench measures a harder problem than "can an agent do useful work" and Zhang is making a broader claim than his own data supports.

Even so, a 40% failure rate on real business tasks is the kind of number that belongs in every pitch deck claiming an agent handles a workflow, and right now almost none of them include it.

ShareXLinkedInEmail

Key Sources

2 sources

Reported by Fortune · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.