Quick answer: Qwen3 235B A22B Instruct 2507 is the July 2025 checkpoint of Qwen3 235B A22B, scoring 83% MMLU-Pro, 79.2% Arena Hard v2, 70.9% BFCL-v3, and 41.8% ARC-AGI. Apache 2.0.
Where Qwen3 235B A22B Instruct 2507 leads
Where it lags
Best for: July 2025 improvements over the base Qwen3 235B; Arena Hard-heavy instruction pipelines; open-source BFCL function calling.
Qwen3 235B A22B Instruct 2507 is the July 2025 (2507) updated checkpoint of the Qwen3 235B A22B instruct model. The "2507" suffix indicates a July 2025 re-release with instruction tuning improvements over the original May 2025 release.
With 79.2% Arena Hard v2 and 70.9% BFCL-v3, this checkpoint shows better instruction following than the base Qwen3 235B A22B release.
| Field | Value |
|---|---|
| Organization | Alibaba |
| License | Apache 2.0 |
| HuggingFace | Qwen/Qwen3-235B-A22B-Instruct-2507 |
| Release date | July 2025 |
| Parameters | 235B total / 22B active (MoE) |
| Modality | Text only |
Open weights under Apache 2.0 — self-host at no cost.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| MMLU-Pro | 83% | Benchgen evaluation | 2025-07 |
| Arena Hard v2 | 79.2% | Benchgen evaluation | 2025-07 |
| BFCL-v3 | 70.9% | Benchgen evaluation | 2025-07 |
| ARC-AGI | 41.8% | Benchgen evaluation | 2025-07 |
| Model | MMLU-Pro | Arena Hard | BFCL-v3 | License |
|---|---|---|---|---|
| Qwen3 235B A22B Instruct 2507 | 83% | 79.2% | 70.9% | Apache 2.0 |
| Qwen3 235B A22B | 68.2% | — | — | Apache 2.0 |
| Qwen3.6 Plus | — | — | — | Proprietary |
Qwen3 235B 2507 vs the base Qwen3 235B: MMLU-Pro 83% vs 68.2% — substantial July 2025 improvement. Use the 2507 checkpoint for better instruction following.
Specs from Alibaba's Qwen3 235B A22B Instruct 2507 release (July 2025) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.