Benchgen
Models/alibaba/

Qwen3.8 Max

DraftPublic

Model Details

Qwen3.8 Max

Organization Params License Modality Released

Quick answer: Qwen3.8 Max (officially "Qwen3.8-Max") is Alibaba's August 2026 flagship — a 2.4T total / 95B active parameter model built on the Qwen 3.5 architecture, and the first Qwen-Max-class model to open-source its weights. It scores 86.6% on Terminal Bench 2.1, 92.6% on GPQA Diamond, 82.8% on IFBench, and 43.6% on Humanity's Last Exam. Available now via QwenCloud API; open weights release the week of August 10, 2026.

At a Glance

Where Qwen3.8 Max leads

  • 86.6% Terminal Bench 2.1 — best-in-class long-horizon terminal agent performance
  • 92.6% GPQA Diamond — frontier-tier graduate-level science reasoning
  • 82.8% IFBench — strongest instruction-following score among compared models
  • First open-weight model at Qwen-Max scale (2.4T params, 95B active)
  • 1M token context window (per Codex model catalog config)

Where it lags

  • 43.6% HLE — behind Claude Opus 4.8 (53.3%) on frontier-exam reasoning
  • 67.7% SWE-bench Pro — behind Claude Opus 4.8 (80.0%) on complex software engineering
  • Open weights not yet available at launch (releasing ~1 week after announcement)

Best for: long-horizon terminal/coding agents, instruction-following-heavy workflows, and teams wanting an open-weight Qwen-Max-class model for self-hosting once weights ship.

What Qwen3.8 Max Is

Qwen3.8-Max is Alibaba's most capable model to date, built on the architectural foundation of Qwen 3.5 and scaled to 2.4 trillion total parameters (95B active). It is the first time Alibaba has open-sourced the weights of a Qwen-Max-class flagship — weights for Hugging Face and ModelScope are slated to ship about a week after the initial API launch. The model supports configurable reasoning_effort levels (xhigh, medium, low) and ships with preserve_thinking enabled by default.

Alibaba's release emphasizes long-horizon autonomous work over single-turn benchmark chasing: showcased case studies include a 10+ day autonomous coding run that self-evolved its own harness, a from-scratch reproduction (and improvement, +2.7pt AIME24) of a published data-selection research paper, and a win in a real 526-team online ML competition (WWW2025 Multimodal Dialogue Intent Recognition Challenge, placing ahead of 458 human teams). On Benchgen's read, Qwen3.8 Max's strongest, most verifiable public-benchmark gains are in agentic coding (Terminal Bench 2.1, FrontierSWE) and instruction-following (IFBench), while frontier-exam reasoning (HLE) and complex SWE-bench-style repo work still trail Claude Opus 4.8.

Specifications

FieldValue
OrganizationAlibaba
Parameters2.4T total, 95B active (MoE)
ArchitectureQwen 3.5 foundation
LicenseApache 2.0 (open weights, releasing ~1 week post-announcement)
Release dateAugust 2026
ModalityMultimodal (text, image; video/document understanding)
Context window1,000,000 tokens

Pricing

Available now via QwenCloud (model slug qwen3.8-max), with OpenAI-compatible chat completions/responses APIs and an Anthropic-compatible API. Refer to QwenCloud/DashScope pricing for current rates. Open weights (Apache 2.0) will allow self-hosting once released.

Public Benchmark Scores

BenchmarkScoreSourceDate
Terminal Bench 2.186.6%Qwen3.8-Max launch blog2026-08
SWE-bench Pro67.7%Qwen3.8-Max launch blog2026-08
DeepSWE 1.156.6%Qwen3.8-Max launch blog2026-08
FrontierSWE73.5%Qwen3.8-Max launch blog2026-08
MLS-Bench-Lite41.0%Qwen3.8-Max launch blog2026-08
JobBench53.4%Qwen3.8-Max launch blog2026-08
Agents' Last Exam52.4 (score) / 27.0% (pass)Qwen3.8-Max launch blog2026-08
Automation-Bench27.3%Qwen3.8-Max launch blog2026-08
Toolathlon Verified72.5%Qwen3.8-Max launch blog2026-08
GPQA Diamond92.6%Qwen3.8-Max launch blog2026-08
Humanity's Last Exam43.6%Qwen3.8-Max launch blog2026-08
IFBench82.8%Qwen3.8-Max launch blog2026-08
MRCR v292.9% (256K, 8-needle)Qwen3.8-Max launch blog2026-08

Scores are self-reported by Alibaba's Qwen Team in the official Qwen3.8-Max launch post, evaluated in-house unless otherwise footnoted (e.g. Terminal Bench 2.1 and FrontierSWE comparison scores for other models are sourced from Artificial Analysis / official leaderboards as of Aug 3, 2026).

Notable In-House / Qualitative Results

Alibaba also reports results on several proprietary, internal evaluations not yet tracked as standalone Benchgen benchmarks:

  • Autonomous chip design — Qwen3.8-Max independently ran a full RTL-to-layout silicon design flow for a GCD/RSA crypto accelerator, reducing gate count from 8,298 to 678 gates (81% physical die-area reduction) over ~500 autonomous turns.
  • E-Commerce Bench (365-day e-commerce operations simulation, internal) — reached a final balance of ¥416,252 (4.16x return), beating second-place GLM 5.2 by 38% and its own prior generation Qwen3.7-Max by 152%.
  • WWW2025 Multimodal Dialogue Intent Recognition Challenge — placed ahead of 458 of 526 competing human teams (0.853 accuracy) in a 24-hour live online competition.
  • QwenSWEBench / QwenQoderBench / QwenReactBench / QwenSVGBench / CoWorkBench / WorkSpaceBench / RecreationBench — Alibaba in-house agentic-coding and cowork benchmarks; not independently reproducible, so not listed as Benchgen leaderboard entries.

Qwen3.8 Max vs Alternatives

ModelTerminal Bench 2.1SWE-bench ProGPQA DiamondHLELicense
Qwen3.8 Max86.6%67.7%92.6%43.6%Apache 2.0
Claude Opus 4.884.6%80.0%92.6%53.3%Proprietary
Claude Fable 588.8%64.6%94.1%47.2%Proprietary
Qwen3.7 Max74.5%60.6%92.4%41.4%Proprietary

Qwen3.8 Max leads its own prior generation (Qwen3.7 Max) across every compared metric, and edges out Claude Opus 4.8 on Terminal Bench 2.1 (86.6% vs 84.6%) and matches it on GPQA Diamond (92.6%), while trailing on SWE-bench Pro and HLE. Claude Fable 5 remains the strongest on raw Terminal Bench 2.1 and GPQA Diamond among this set.

FAQ

Is Qwen3.8 Max open source? Yes — it's the first Qwen-Max-class model with open weights. The model launched via API on August 3, 2026, with Hugging Face/ModelScope weights following about a week later.

How big is Qwen3.8 Max? 2.4 trillion total parameters with 95 billion active parameters (mixture-of-experts), built on the Qwen 3.5 architecture.

What is Qwen3.8 Max's context window? 1,000,000 tokens, per Alibaba's published Codex model-catalog configuration.

How does Qwen3.8 Max compare to Claude Opus 4.8? Qwen3.8 Max leads on Terminal Bench 2.1 (86.6% vs 84.6%) and ties on GPQA Diamond (92.6%), but trails on SWE-bench Pro (67.7% vs 80.0%) and HLE (43.6% vs 53.3%).

Where can I access Qwen3.8 Max? Via QwenCloud now (model slug qwen3.8-max), with OpenAI- and Anthropic-compatible APIs; open weights follow roughly a week after launch.


Benchmark scores sourced from Alibaba's official Qwen3.8-Max launch announcement (qwen.ai, Aug 3, 2026). Comparison-model scores for Claude Opus 4.8, Claude Fable 5, and Qwen3.7 Max are as reported in that same table.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.