Analysis
Alibaba released Qwen3.8-Max this week with a specific, headline-grabbing claim: on OSWorld-Verified, a benchmark that measures how well an AI agent can actually operate a computer's operating system and applications, Qwen3.8-Max scores 86.1, ahead of OpenAI's GPT-5.6 Sol Max at 83.2 and Anthropic's Fable 5 at 85.0. Agentic computer-use is one of the most closely watched capability categories right now, given the enterprise push toward AI agents that can operate software directly rather than just generate text.
The standard caveat applies: these are vendor-reported numbers from Alibaba's own launch materials, not independently verified results, and benchmark scores are highly sensitive to the specific evaluation harness used -- a pattern that's made nearly every frontier-lab benchmark claim this cycle worth treating with some skepticism until third-party verification catches up.
“The broader picture is more nuanced than the headline claim suggests.”
The broader picture is more nuanced than the headline claim suggests. Qwen reports leading results on PaperBench and strong marks on TerminalBench and long-video understanding, but Fable 5 remains ahead on several software-engineering benchmarks specifically -- including a gap of more than 13 points over Qwen3.8-Max on DeepSWE 1.1 (70.0 versus 56.6). Different models are winning different slices of the evaluation landscape depending on task type.
This continues an accelerating pattern where Chinese open-weight labs -- Qwen, DeepSeek, MiniMax, Kimi -- routinely claim parity or superiority against OpenAI and Anthropic's latest frontier releases on at least some benchmark categories, even as questions about safety disclosure and red-teaming rigor (see the open-weight safety-gap story elsewhere in this issue) remain largely unresolved for the same models.
What to watch: whether independent third-party evaluators replicate Qwen3.8-Max's OSWorld-Verified score once the model is more widely tested outside Alibaba's own benchmarking environment, and whether agentic computer-use becomes the next capability category where open-weight models genuinely lead rather than merely claim parity.