VC
Value Add VC
⚡HomePulse⚡Helpful Apps📝Blog🤝Partner
Illustration for: Alibaba's Qwen3.8-Max Claims It Beats GPT-5.6, Fable 5
Value Add VC/Pulse/AI

Alibaba's Qwen3.8-Max Claims It Beats GPT-5.6, Fable 5

Alibaba's Qwen3.8-Max reports 86.1 on the OSWorld-Verified computer-use benchmark versus 83.2 for GPT-5.6 Sol Max and 85.0 for Fable 5, though the picture is murkier once other benchmarks are included.

TC
Trace Cohen
Early-stage VC & angel · Founder, New York Venture Partners
August 3, 2026
1 min read
ShareXLinkedInEmail

THE RUNDOWN

1

Qwen3.8-Max reports 86.1 on OSWorld-Verified, a benchmark measuring how well AI agents can operate a computer OS and its applications, ahead of GPT-5.6 Sol Max's 83.2 and Fable 5's 85.0

2

The numbers are vendor-reported from Alibaba's own launch materials, not independently verified -- a standard caveat that applies to nearly every frontier-model benchmark claim this cycle

3

The picture is more mixed elsewhere: Fable 5 leads several software-engineering benchmarks, including a 13+ point gap over Qwen3.8-Max on DeepSWE 1.1 (70.0 vs 56.6)

4

It's another entry in the fast-accelerating open-weight vs. closed-lab benchmark race, with Qwen, DeepSeek, MiniMax and Kimi all now routinely claiming parity or superiority on specific benchmark slices against OpenAI and Anthropic's frontier models

TC

The VC Read · Trace's Take

Trace Cohen

Vendor-reported benchmarks are marketing until a third party replicates them -- I'd treat the 86.1 number as a claim, not a fact, until it holds up outside Alibaba's own harness. The more interesting signal is that agentic computer-use is now contested territory at all; a year ago nobody was racing on this specific capability.

AI Valuations Tracker →

Analysis

Alibaba released Qwen3.8-Max this week with a specific, headline-grabbing claim: on OSWorld-Verified, a benchmark that measures how well an AI agent can actually operate a computer's operating system and applications, Qwen3.8-Max scores 86.1, ahead of OpenAI's GPT-5.6 Sol Max at 83.2 and Anthropic's Fable 5 at 85.0. Agentic computer-use is one of the most closely watched capability categories right now, given the enterprise push toward AI agents that can operate software directly rather than just generate text.

The standard caveat applies: these are vendor-reported numbers from Alibaba's own launch materials, not independently verified results, and benchmark scores are highly sensitive to the specific evaluation harness used -- a pattern that's made nearly every frontier-lab benchmark claim this cycle worth treating with some skepticism until third-party verification catches up.

“The broader picture is more nuanced than the headline claim suggests.”

The broader picture is more nuanced than the headline claim suggests. Qwen reports leading results on PaperBench and strong marks on TerminalBench and long-video understanding, but Fable 5 remains ahead on several software-engineering benchmarks specifically -- including a gap of more than 13 points over Qwen3.8-Max on DeepSWE 1.1 (70.0 versus 56.6). Different models are winning different slices of the evaluation landscape depending on task type.

This continues an accelerating pattern where Chinese open-weight labs -- Qwen, DeepSeek, MiniMax, Kimi -- routinely claim parity or superiority against OpenAI and Anthropic's latest frontier releases on at least some benchmark categories, even as questions about safety disclosure and red-teaming rigor (see the open-weight safety-gap story elsewhere in this issue) remain largely unresolved for the same models.

What to watch: whether independent third-party evaluators replicate Qwen3.8-Max's OSWorld-Verified score once the model is more widely tested outside Alibaba's own benchmarking environment, and whether agentic computer-use becomes the next capability category where open-weight models genuinely lead rather than merely claim parity.

ShareXLinkedInEmail

More on

Anthropic →Alibaba →

Reported by VentureBeat · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.

Read Next

AI· Aug 9, 2026

Physical AI's Biggest Week Yet: $21 Billion

Illustration for: Physical AI's Biggest Week Yet: $21 Billion
AI

Physical AI's Biggest Week Yet: $21 Billion

Six deals in seven days — Lumilens, Hadrian, Terafab, Valar Atomics, Base Power and K2 Space — pushed more than $21 billion into reactors, factories, satellites and chips, not a single model release among them.

AI· Aug 6, 2026

Claude Code Adds Self-Hosted Session Environments

Illustration for: Claude Code Adds Self-Hosted Session Environments
AI

Claude Code Adds Self-Hosted Session Environments

Anthropic opened a public beta letting Claude Code sessions run on a customer's own infrastructure instead of Anthropic's cloud, aimed at teams whose compliance or network requirements ruled out the hosted version.

AI· Aug 7, 2026

Why the AI Labs Just Rewired Their Org Charts

Illustration for: Why the AI Labs Just Rewired Their Org Charts
AI

Why the AI Labs Just Rewired Their Org Charts

Hassabis moving to chair, Jeff Dean's exit, and Anthropic's new chip team all landed in one week -- a trace take on what it means that frontier labs are restructuring around infrastructure, not research.

@Trace_Cohen·t@nyvp.com