VC
Value Add VC
⚡HomePulse⚡Helpful Apps📝Blog🤝Partner
Illustration for: Alibaba's Qwen3.8-Max Claims It Beats GPT-5.6, Fable 5
Value Add VC/Pulse/AI

Alibaba's Qwen3.8-Max Claims It Beats GPT-5.6, Fable 5

Alibaba's Qwen3.8-Max reports 86.1 on the OSWorld-Verified computer-use benchmark versus 83.2 for GPT-5.6 Sol Max and 85.0 for Fable 5, though the picture is murkier once other benchmarks are included.

TC
Trace Cohen
Early-stage VC & angel · Founder, New York Venture Partners
August 3, 2026
1 min read
ShareXLinkedInEmail

THE RUNDOWN

1

Qwen3.8-Max reports 86.1 on OSWorld-Verified, a benchmark measuring how well AI agents can operate a computer OS and its applications, ahead of GPT-5.6 Sol Max's 83.2 and Fable 5's 85.0

2

The numbers are vendor-reported from Alibaba's own launch materials, not independently verified -- a standard caveat that applies to nearly every frontier-model benchmark claim this cycle

3

The picture is more mixed elsewhere: Fable 5 leads several software-engineering benchmarks, including a 13+ point gap over Qwen3.8-Max on DeepSWE 1.1 (70.0 vs 56.6)

4

It's another entry in the fast-accelerating open-weight vs. closed-lab benchmark race, with Qwen, DeepSeek, MiniMax and Kimi all now routinely claiming parity or superiority on specific benchmark slices against OpenAI and Anthropic's frontier models

TC

The VC Read · Trace's Take

Trace Cohen

Vendor-reported benchmarks are marketing until a third party replicates them -- I'd treat the 86.1 number as a claim, not a fact, until it holds up outside Alibaba's own harness. The more interesting signal is that agentic computer-use is now contested territory at all; a year ago nobody was racing on this specific capability.

AI Valuations Tracker →

Analysis

Alibaba released Qwen3.8-Max this week with a specific, headline-grabbing claim: on OSWorld-Verified, a benchmark that measures how well an AI agent can actually operate a computer's operating system and applications, Qwen3.8-Max scores 86.1, ahead of OpenAI's GPT-5.6 Sol Max at 83.2 and Anthropic's Fable 5 at 85.0. Agentic computer-use is one of the most closely watched capability categories right now, given the enterprise push toward AI agents that can operate software directly rather than just generate text.

The standard caveat applies: these are vendor-reported numbers from Alibaba's own launch materials, not independently verified results, and benchmark scores are highly sensitive to the specific evaluation harness used -- a pattern that's made nearly every frontier-lab benchmark claim this cycle worth treating with some skepticism until third-party verification catches up.

“The broader picture is more nuanced than the headline claim suggests.”

The broader picture is more nuanced than the headline claim suggests. Qwen reports leading results on PaperBench and strong marks on TerminalBench and long-video understanding, but Fable 5 remains ahead on several software-engineering benchmarks specifically -- including a gap of more than 13 points over Qwen3.8-Max on DeepSWE 1.1 (70.0 versus 56.6). Different models are winning different slices of the evaluation landscape depending on task type.

This continues an accelerating pattern where Chinese open-weight labs -- Qwen, DeepSeek, MiniMax, Kimi -- routinely claim parity or superiority against OpenAI and Anthropic's latest frontier releases on at least some benchmark categories, even as questions about safety disclosure and red-teaming rigor (see the open-weight safety-gap story elsewhere in this issue) remain largely unresolved for the same models.

What to watch: whether independent third-party evaluators replicate Qwen3.8-Max's OSWorld-Verified score once the model is more widely tested outside Alibaba's own benchmarking environment, and whether agentic computer-use becomes the next capability category where open-weight models genuinely lead rather than merely claim parity.

ShareXLinkedInEmail

More on

Anthropic →Alibaba →

Analysis and editorial commentary by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.

Read Next

AI· Aug 4, 2026

Open-Weight AI Closes Gap, Not Safety Gap

Illustration for: Open-Weight AI Closes Gap, Not Safety Gap
AI

Open-Weight AI Closes Gap, Not Safety Gap

Open-weight AI models are approaching frontier-lab performance on many benchmarks, but researchers say safety tooling and guardrails for open models still lag well behind what closed labs have built.

AI· Aug 4, 2026

AI Coding Agents Are Blowing Through Startup Budgets

Illustration for: AI Coding Agents Are Blowing Through Startup Budgets
AI

AI Coding Agents Are Blowing Through Startup Budgets

Companies like Replit, Kilo Code and Symbotic say AI coding agent usage is scaling costs far faster than teams expected, forcing new usage-monitoring and budgeting practices around agent-driven development.

AI· Aug 5, 2026

Google Assistant Dies September 4, Gemini Takes Over

Illustration for: Google Assistant Dies September 4, Gemini Takes Over
AI

Google Assistant Dies September 4, Gemini Takes Over

Google will begin removing Assistant from Android phones, tablets, Wear OS and Android Auto on September 4, replacing it entirely with Gemini with no option to switch back.

@Trace_Cohen·t@nyvp.com