Skip to content

Alibaba launches Qwen3.8-Max globally: 2.4T MoE edges Fable 5 on computer-use benchmark, open weights due August 10

· by Pondero Newsdesk

The short version

Alibaba opened global API access to Qwen3.8-Max on August 3, posting 86.1 on OSWorld-Verified against Fable 5's 85.0. The 2.4-trillion-parameter MoE model starts at $2 per million input tokens; open weights and the QwenWork enterprise agent platform arrive within the week.

Alibaba launches Qwen3.8-Max globally: 2.4T MoE edges Fable 5 on computer-use benchmark, open weights due August 10

OSWorld-Verified, the benchmark that now separates top-tier frontier models by a few percentage points, put Qwen3.8-Max at 86.1, ahead of Claude Fable 5 at 85.0 and GPT-5.6 Sol Max at 83.2. Alibaba opened global API access on August 3 for developers via its Cloud Model Studio, with open weights and a new enterprise agent platform set to follow within days.

What Alibaba shipped

Qwen3.8-Max is a mixture-of-experts model with 2.4 trillion total parameters and 95 billion active at inference, per Alibaba Cloud's announcement. It accepts text, image, and video input with a 1-million-token context window, large enough to process roughly 200 pages of text or about 100 hours of video per request. Function calling, structured outputs, and fine-tuning are all supported.

API pricing on Alibaba Cloud Model Studio is $2.00 per million input tokens and $6.00 per million output tokens, cached input at $0.25 per million tokens, per MarkTechPost. At those rates, Qwen3.8-Max comes in below the stated price of comparable frontier models from Anthropic and OpenAI.

Alibaba also opened QwenWork to public beta: an all-in-one workplace agent interface covering writing, coding, and research tasks for individuals and teams. To demonstrate the model, Alibaba ran a 16-day autonomous coding project and a 500-step chip design optimization, though both are vendor-run showcases rather than independent evaluations.

Benchmark context

Alibaba's headline claim is the 86.1 score on OSWorld-Verified, which tests agentic desktop computer use by scoring task completion on real GUI environments without human intervention. Three models cluster within about three points at the top: Fable 5 at 85.0, GPT-5.6 Sol Max at 83.2, and Qwen3.8-Max at 86.1, per OfficeChai.

Alibaba also published 93.0 on PaperBench, 92.6 on GPQA Diamond, and 56.6 on DeepSWE 1.1. DeepSWE 1.1 improved from 21.6 on the prior generation, a substantial gain on that coding benchmark. Multiple outlets reporting from the announcement put active parameters at 95 billion; Alibaba's own documentation had not explicitly confirmed that figure as of publication.

Why it matters

With just 1.1 points separating Qwen3.8-Max from Fable 5 on OSWorld-Verified, operators choosing an agentic provider should run their own task suite before selecting a vendor on this benchmark alone. Cost is the clearer differentiator: $2 per million input tokens is below what comparable frontier models charge, which creates a cost-efficiency case for teams already routing data through Alibaba's cloud.

Open weights arriving around August 10 matter most to teams that cannot send data to a Chinese cloud provider. A permissive commercial license would let regulated-industry operators deploy a 2.4T MoE on their own infrastructure, which few models at this scale support. License terms had not been confirmed as of publication.

TechTimes flagged that QwenWork routes enterprise workflows through Alibaba's infrastructure, raising data-residency considerations for teams subject to US or EU compliance rules, per TechTimes. A direct vendor conversation before any QwenWork procurement is warranted.

What to watch next

Alibaba has not confirmed a specific open-weights date beyond "next week" from August 3. License terms will determine whether self-hosted deployment is commercially viable and in which jurisdictions. An August OSWorld-Verified leaderboard update will confirm whether the 86.1 score holds as other labs submit new results.

Sources