Benchgen
Models/alibaba/

Qwen3-235B-A22B-Thinking-2507

DraftPublic

Model Details

Qwen3 235B A22B Thinking 2507

Organization License Released Thinking

Quick answer: Qwen3 235B A22B Thinking 2507 is the July 2025 checkpoint of the Qwen3 235B A22B thinking model, scoring 84.4% MMLU-Pro, 79.7% Arena Hard v2, and 71.9% BFCL-v3. Apache 2.0.

At a Glance

Where Qwen3 235B A22B Thinking 2507 leads

  • 84.4% MMLU-Pro — improved over Instruct 2507 (83%)
  • 71.9% BFCL-v3 — strong function calling
  • 79.7% Arena Hard — competitive
  • Apache 2.0 — fully open

Where it lags

  • 235B total params: large hosting requirement

Best for: Maximum Qwen3 235B thinking quality; MMLU-Pro + BFCL combined pipelines; July 2025 thinking checkpoint improvements.

What Qwen3 235B A22B Thinking 2507 Is

The thinking (chain-of-thought) July 2025 checkpoint of Qwen3 235B A22B. The "Thinking" variant uses extended chain-of-thought, improving MMLU-Pro (84.4% vs 83%) and BFCL (71.9% vs 70.9%) over the instruct variant.

Specifications

FieldValue
OrganizationAlibaba
LicenseApache 2.0
HuggingFaceQwen/Qwen3-235B-A22B-Thinking-2507
Release dateJuly 2025
Parameters235B total / 22B active (MoE)
ModalityText only

Pricing

Open weights under Apache 2.0 — self-host at no cost.

Public Benchmark Scores

BenchmarkScoreSourceDate
MMLU-Pro84.4%Benchgen evaluation2025-07
Arena Hard v279.7%Benchgen evaluation2025-07
BFCL-v371.9%Benchgen evaluation2025-07

Qwen3 235B A22B Thinking 2507 vs Alternatives

ModelMMLU-ProArena HardBFCL-v3Type
Qwen3 235B A22B Thinking 250784.4%79.7%71.9%Thinking
Qwen3 235B A22B Instruct 250783%79.2%70.9%Instruct

Thinking 2507 vs Instruct 2507: +1.4% MMLU-Pro, +0.5% Arena Hard, +1.0% BFCL. Small improvements from thinking chain-of-thought.

Frequently Asked Questions

What is Qwen3 235B A22B Thinking 2507? Alibaba's July 2025 thinking MoE checkpoint scoring 84.4% MMLU-Pro, 79.7% Arena Hard. Apache 2.0.

Specs from Alibaba's Qwen3 235B A22B Thinking 2507 release (July 2025) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.