Benchgen
Models/alibaba/

Qwen3 235B A22B

DraftPublic

Model Details

Qwen3 235B A22B

Organization License Released MoE Params

Quick answer: Qwen3 235B A22B is Alibaba's April 2025 flagship MoE model, scoring 95.6% Arena Hard, 70.8% BFCL, 65.9% LiveCodeBench, 94.4% GSM8K, and 61.8% Aider. Apache 2.0 with 235B total parameters (22B active). Alibaba's strongest open-weight model.

At a Glance

Where Qwen3 235B A22B leads

  • 95.6% Arena Hard — near-frontier instruction quality; above DeepSeek-V3 (76.2%)
  • 65.9% LiveCodeBench — competitive coding with frontier closed models
  • 70.8% BFCL — strong function calling
  • 94.4% GSM8K — excellent math
  • Apache 2.0 — fully open, including fine-tuning and redistribution
  • MoE efficiency: 22B active params at inference vs 235B total

Where it lags

  • 68.2% MMLU-Pro — moderate academic knowledge
  • 71.8% MATH — below strong reasoning models
  • 17.6% AetherCode — weak on this agent benchmark
  • Requires large GPU cluster for full deployment

Best for: Open-source frontier instruction quality; tool-use pipelines; teams needing Apache 2.0 near-frontier models on shared infrastructure.

What Qwen3 235B A22B Is

Qwen3 235B A22B is Alibaba's April 2025 flagship in the Qwen3 generation — succeeding Qwen2.5. The "235B A22B" notation means 235 billion total parameters with 22 billion active at each token inference pass (MoE routing). This gives DeepSeek-level per-token compute cost while maintaining the capacity of a 235B model.

The 95.6% Arena Hard score is exceptional for an open-weight model — competitive with proprietary models like Claude Opus-class and GPT-4-class systems from the same period. Combined with Apache 2.0, this is a strong option for teams needing frontier-class instruction quality without vendor lock-in.

Specifications

FieldValue
OrganizationAlibaba
LicenseApache 2.0
HuggingFaceQwen/Qwen3-235B-A22B
Release dateApril 29, 2025
Parameters235B total / 22B active (MoE)
ModalityText only
Context window128K tokens

Pricing

Open weights under Apache 2.0 — self-host. Also available via Alibaba Cloud (DashScope) and major providers.

Public Benchmark Scores

BenchmarkScoreSourceDate
Arena Hard95.6%Benchgen evaluation2025-04
GSM8K94.4%Benchgen evaluation2025-04
BFCL70.8%Benchgen evaluation2025-04
LiveCodeBench65.9%Benchgen evaluation2025-04
MATH71.8%Benchgen evaluation2025-04
MMLU-Pro68.2%Benchgen evaluation2025-04
Aider61.8%Benchgen evaluation2025-04

Qwen3 235B A22B vs Alternatives

ModelArena HardLiveCodeBenchLicenseArchitecture
Qwen3 235B A22B95.6%65.9%Apache 2.0MoE 235B/22B
DeepSeek-V376.2%MITMoE
Qwen3 30B A3B91%62.6%Apache 2.0MoE 30B/3B
Llama 3.3 70B InstructLlama 3.3Dense

Qwen3 235B A22B leads open-weight models on Arena Hard (95.6%). For lower compute: Qwen3 30B A3B (91% Arena Hard, 3B active params). For general open-weight: DeepSeek-V3.

Frequently Asked Questions

What is Qwen3 235B A22B? Alibaba's April 2025 flagship MoE model: 235B total parameters, 22B active at inference. 95.6% Arena Hard, 65.9% LiveCodeBench, 70.8% BFCL. Apache 2.0.
What does "235B A22B" mean in Qwen3? 235B = total parameter count. A22B = 22B parameters active per token inference pass. Mixture-of-Experts routing selects which expert layers to activate, making inference cost similar to a 22B dense model while retaining 235B model capacity.

Specs from Alibaba's Qwen3 235B A22B release (April 2025) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.