Quick answer: Qwen3 235B A22B Thinking 2507 is the July 2025 checkpoint of the Qwen3 235B A22B thinking model, scoring 84.4% MMLU-Pro, 79.7% Arena Hard v2, and 71.9% BFCL-v3. Apache 2.0.
Where Qwen3 235B A22B Thinking 2507 leads
Where it lags
Best for: Maximum Qwen3 235B thinking quality; MMLU-Pro + BFCL combined pipelines; July 2025 thinking checkpoint improvements.
The thinking (chain-of-thought) July 2025 checkpoint of Qwen3 235B A22B. The "Thinking" variant uses extended chain-of-thought, improving MMLU-Pro (84.4% vs 83%) and BFCL (71.9% vs 70.9%) over the instruct variant.
| Field | Value |
|---|---|
| Organization | Alibaba |
| License | Apache 2.0 |
| HuggingFace | Qwen/Qwen3-235B-A22B-Thinking-2507 |
| Release date | July 2025 |
| Parameters | 235B total / 22B active (MoE) |
| Modality | Text only |
Open weights under Apache 2.0 — self-host at no cost.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| MMLU-Pro | 84.4% | Benchgen evaluation | 2025-07 |
| Arena Hard v2 | 79.7% | Benchgen evaluation | 2025-07 |
| BFCL-v3 | 71.9% | Benchgen evaluation | 2025-07 |
| Model | MMLU-Pro | Arena Hard | BFCL-v3 | Type |
|---|---|---|---|---|
| Qwen3 235B A22B Thinking 2507 | 84.4% | 79.7% | 71.9% | Thinking |
| Qwen3 235B A22B Instruct 2507 | 83% | 79.2% | 70.9% | Instruct |
Thinking 2507 vs Instruct 2507: +1.4% MMLU-Pro, +0.5% Arena Hard, +1.0% BFCL. Small improvements from thinking chain-of-thought.
Specs from Alibaba's Qwen3 235B A22B Thinking 2507 release (July 2025) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.