Quick answer: Qwen3 235B A22B is Alibaba's April 2025 flagship MoE model, scoring 95.6% Arena Hard, 70.8% BFCL, 65.9% LiveCodeBench, 94.4% GSM8K, and 61.8% Aider. Apache 2.0 with 235B total parameters (22B active). Alibaba's strongest open-weight model.
Where Qwen3 235B A22B leads
Where it lags
Best for: Open-source frontier instruction quality; tool-use pipelines; teams needing Apache 2.0 near-frontier models on shared infrastructure.
Qwen3 235B A22B is Alibaba's April 2025 flagship in the Qwen3 generation — succeeding Qwen2.5. The "235B A22B" notation means 235 billion total parameters with 22 billion active at each token inference pass (MoE routing). This gives DeepSeek-level per-token compute cost while maintaining the capacity of a 235B model.
The 95.6% Arena Hard score is exceptional for an open-weight model — competitive with proprietary models like Claude Opus-class and GPT-4-class systems from the same period. Combined with Apache 2.0, this is a strong option for teams needing frontier-class instruction quality without vendor lock-in.
| Field | Value |
|---|---|
| Organization | Alibaba |
| License | Apache 2.0 |
| HuggingFace | Qwen/Qwen3-235B-A22B |
| Release date | April 29, 2025 |
| Parameters | 235B total / 22B active (MoE) |
| Modality | Text only |
| Context window | 128K tokens |
Open weights under Apache 2.0 — self-host. Also available via Alibaba Cloud (DashScope) and major providers.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| Arena Hard | 95.6% | Benchgen evaluation | 2025-04 |
| GSM8K | 94.4% | Benchgen evaluation | 2025-04 |
| BFCL | 70.8% | Benchgen evaluation | 2025-04 |
| LiveCodeBench | 65.9% | Benchgen evaluation | 2025-04 |
| MATH | 71.8% | Benchgen evaluation | 2025-04 |
| MMLU-Pro | 68.2% | Benchgen evaluation | 2025-04 |
| Aider | 61.8% | Benchgen evaluation | 2025-04 |
| Model | Arena Hard | LiveCodeBench | License | Architecture |
|---|---|---|---|---|
| Qwen3 235B A22B | 95.6% | 65.9% | Apache 2.0 | MoE 235B/22B |
| DeepSeek-V3 | 76.2% | — | MIT | MoE |
| Qwen3 30B A3B | 91% | 62.6% | Apache 2.0 | MoE 30B/3B |
| Llama 3.3 70B Instruct | — | — | Llama 3.3 | Dense |
Qwen3 235B A22B leads open-weight models on Arena Hard (95.6%). For lower compute: Qwen3 30B A3B (91% Arena Hard, 3B active params). For general open-weight: DeepSeek-V3.
Specs from Alibaba's Qwen3 235B A22B release (April 2025) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.