Benchgen
Models/alibaba/

Qwen3.6-27B

DraftPublic

Model Details

Qwen3.6 27B

Organization License Released Params

Quick answer: Qwen3.6 27B is Alibaba's December 2025 dense 27B model scoring 94.1% AIME 2026, 87.8% GPQA Diamond, 86.2% MMLU-Pro, 77.2% SWE-Bench Verified, and 75.8% MMMU-Pro. Apache 2.0.

At a Glance

Where Qwen3.6 27B leads

  • 94.1% AIME 2026 — exceptional competition math
  • 87.8% GPQA Diamond — strong graduate science
  • 86.2% MMLU-Pro — excellent academic breadth
  • 77.2% SWE-Bench Verified — strong software engineering
  • 75.8% MMMU-Pro — solid multimodal
  • Apache 2.0 — fully open

Where it lags

  • Dense 27B: higher inference cost than the 35B A3B MoE sibling
  • No BrowseComp/HLE scores available

Best for: Open-source AIME competition math; combined GPQA + SWE-Bench Verified at 27B scale; cost-balanced alternative to Qwen3.6 35B A3B.

What Qwen3.6 27B Is

Qwen3.6 27B is the dense 27B variant in the Qwen3.6 generation (December 2025). Unlike the 35B A3B MoE sibling (3B active), this is a full 27B dense model — slightly higher inference cost but simpler deployment.

The 94.1% AIME 2026 is the standout metric — one of the highest AIME scores for any 27B model. Combined with 77.2% SWE-Bench Verified, this positions Qwen3.6 27B as a top open-source model for both math and coding.

Specifications

FieldValue
OrganizationAlibaba
LicenseApache 2.0
HuggingFaceQwen/Qwen3.6-27B
Release dateDecember 2025
Parameters27B (dense)
ModalityText only

Pricing

Open weights under Apache 2.0 — self-host at no cost.

Public Benchmark Scores

BenchmarkScoreSourceDate
AIME 202694.1%Benchgen evaluation2025-12
GPQA Diamond87.8%Benchgen evaluation2025-12
MMLU-Pro86.2%Benchgen evaluation2025-12
SWE-Bench Verified77.2%Benchgen evaluation2025-12
MMMU-Pro75.8%Benchgen evaluation2025-12

Qwen3.6 27B vs Alternatives

ModelAIME 2026GPQA DiamondSWE-Bench VerifiedLicense
Qwen3.6 27B94.1%87.8%77.2%Apache 2.0
Qwen3.6 35B A3B92.7%86.0%Apache 2.0
Qwen3.5 27BApache 2.0
Meta Muse Spark89.5%77.4%Proprietary

Qwen3.6 27B vs Qwen3.6 35B A3B: dense 27B has slightly higher AIME (94.1% vs 92.7%), similar GPQA (87.8% vs 86.0%), and adds SWE-Bench Verified (77.2%). Tradeoff: higher inference cost (27B vs 3B active).

Frequently Asked Questions

What is Qwen3.6 27B? Alibaba's December 2025 dense 27B model scoring 94.1% AIME 2026, 87.8% GPQA Diamond, 77.2% SWE-Bench Verified. Apache 2.0.

Specs from Alibaba's Qwen3.6 27B release (December 2025) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.