Quick answer: Qwen3.6 27B is Alibaba's December 2025 dense 27B model scoring 94.1% AIME 2026, 87.8% GPQA Diamond, 86.2% MMLU-Pro, 77.2% SWE-Bench Verified, and 75.8% MMMU-Pro. Apache 2.0.
Where Qwen3.6 27B leads
Where it lags
Best for: Open-source AIME competition math; combined GPQA + SWE-Bench Verified at 27B scale; cost-balanced alternative to Qwen3.6 35B A3B.
Qwen3.6 27B is the dense 27B variant in the Qwen3.6 generation (December 2025). Unlike the 35B A3B MoE sibling (3B active), this is a full 27B dense model — slightly higher inference cost but simpler deployment.
The 94.1% AIME 2026 is the standout metric — one of the highest AIME scores for any 27B model. Combined with 77.2% SWE-Bench Verified, this positions Qwen3.6 27B as a top open-source model for both math and coding.
| Field | Value |
|---|---|
| Organization | Alibaba |
| License | Apache 2.0 |
| HuggingFace | Qwen/Qwen3.6-27B |
| Release date | December 2025 |
| Parameters | 27B (dense) |
| Modality | Text only |
Open weights under Apache 2.0 — self-host at no cost.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| AIME 2026 | 94.1% | Benchgen evaluation | 2025-12 |
| GPQA Diamond | 87.8% | Benchgen evaluation | 2025-12 |
| MMLU-Pro | 86.2% | Benchgen evaluation | 2025-12 |
| SWE-Bench Verified | 77.2% | Benchgen evaluation | 2025-12 |
| MMMU-Pro | 75.8% | Benchgen evaluation | 2025-12 |
| Model | AIME 2026 | GPQA Diamond | SWE-Bench Verified | License |
|---|---|---|---|---|
| Qwen3.6 27B | 94.1% | 87.8% | 77.2% | Apache 2.0 |
| Qwen3.6 35B A3B | 92.7% | 86.0% | — | Apache 2.0 |
| Qwen3.5 27B | — | — | — | Apache 2.0 |
| Meta Muse Spark | — | 89.5% | 77.4% | Proprietary |
Qwen3.6 27B vs Qwen3.6 35B A3B: dense 27B has slightly higher AIME (94.1% vs 92.7%), similar GPQA (87.8% vs 86.0%), and adds SWE-Bench Verified (77.2%). Tradeoff: higher inference cost (27B vs 3B active).
Specs from Alibaba's Qwen3.6 27B release (December 2025) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.