Analysis
I spent part of this week reading Kuo Zhang's piece in Fortune, and it's a rare thing: an AI executive publishing his own company's benchmark showing his own company's agents failing. Zhang runs Alibaba.com, and his team built CommerceAgentBench, an open-source test of 107 real e-commerce tasks -- sourcing, landed-cost math, multi-leg shipping, after-sales disputes -- drawn from 10 million active small-business users, 1.6 million real conversations and 200,000 execution traces.
The headline number: the strongest frontier model Alibaba tested completed 61.7% of tasks. Put differently, nearly 40% of the time, a task that looked plausible in the transcript came back wrong when someone checked the actual output against the actual system. Zhang's framing is blunt -- agents can talk convincingly about landed costs, currency conversion and shipping routes, and still get the arithmetic wrong in a way that costs a real seller real money.
This matters beyond Alibaba because every major lab is currently selling some version of the opposite story:
“The headline number: the strongest frontier model Alibaba tested completed 61.7% of tasks.”
- OpenAI -- ChatGPT's agent mode is marketed as able to book, buy and file on a user's behalf with minimal supervision.
- Anthropic -- Claude's computer-use and "Cowork" features are pitched as production-ready for multi-step business workflows.
- Google -- Gemini's Project Mariner demos browser agents completing shopping and research tasks autonomously.
None of those companies has published a CommerceAgentBench-style number showing their own agent's real-world task completion rate against a benchmark built from actual paying customers rather than curated demo tasks. Alibaba did, and it isn't flattering.
If I'm underwriting an AI agent startup right now, this changes what I ask for in diligence. A demo video or a cherry-picked eval score tells me almost nothing about whether the agent works on the tenth task, the messy one, the after-sales dispute with three PDFs attached. I want the founder's own failure-mode benchmark -- task categories, sample size, and the percentage that came back wrong -- not the percentage that came back impressive.
Room for disagreement: the 61.7% number is one company's internal benchmark on one company's task distribution, and Alibaba has an incentive to publish a number that makes room for a human-in-the-loop product tier it can upsell. Agent performance has also moved fast -- the frontier model that scored 61.7% today may not be the frontier model in three months, and coding agents already clear far higher completion rates on narrower, more structured tasks than open-ended commerce. It's possible CommerceAgentBench measures a harder problem than "can an agent do useful work" and Zhang is making a broader claim than his own data supports.
Even so, a 40% failure rate on real business tasks is the kind of number that belongs in every pitch deck claiming an agent handles a workflow, and right now almost none of them include it.