Benchgen
Models/alibaba/

Qwen3.8-27B

DraftPublic

Model Details

Qwen3.8-27B

Organization Params License Modality Released

Quick answer: Qwen3.8-27B is Alibaba's compact, deployment-friendly dense model in the Qwen3.8 family — a 27B-parameter native vision-language model (text, image, and hour-scale video understanding) built on the Qwen 3.5 architecture, released under Apache 2.0. It scores 73.0% on Terminal Bench 2.1, 89.2% on GPQA Diamond, 90.3% on LiveCodeBench v6, and 30.8% on Humanity's Last Exam.

At a Glance

Where Qwen3.8-27B leads

  • 90.3% LiveCodeBench v6 — strong competitive-coding performance for a 27B dense model
  • 89.2% GPQA Diamond — near-flagship-tier graduate-level science reasoning at a fraction of the size
  • 84.3% OSWorld-Verified — strong computer-use/agentic desktop control
  • 81.9% AndroidWorld — strong mobile-use agentic control
  • Native vision-language: images, STEM diagrams/documents, and hour-scale video understanding
  • Apache 2.0 — fully open weights, self-hostable on consumer/single-GPU hardware

Where it lags

  • 30.8% HLE — well behind flagship-tier models (e.g. Qwen3.8 Max at 43.6%)
  • 33.4% JobBench — modest on long-horizon professional job-task simulation
  • 42.3% NL2Repo-Bench — modest on repo-level code generation vs flagship models

Best for: compact/self-hosted vision-language agents, computer-use and mobile-use automation, and coding assistants where a smaller footprint matters more than flagship-scale reasoning.

What Qwen3.8-27B Is

Qwen3.8-27B is Alibaba's compact, deployment-friendly dense model in the Qwen3.8 generation, released alongside the flagship Qwen3.8-Max. Built on the same Qwen 3.5 architectural foundation, it's a 27B-parameter causal language model with a native vision encoder — a true vision-language model that understands images (including STEM diagrams and documents) and hour-scale videos out of the box, not a bolted-on multimodal adapter. Thinking mode is on by default (configurable via reasoning_effort: xhigh, medium, low) and reasoning context is retained across turns via preserve_thinking.

The open-weight checkpoint natively supports a 262,144-token context window, extensible to 1,000,000 tokens via YaRN RoPE scaling. Alibaba also plans a hosted version via Qwen Cloud with a 1M-token context by default and built-in tools ("coming soon" per the model card).

Qwen3.8-27B is aimed at agentic and coding workloads that don't require flagship-scale parameters: it scores competitively on LiveCodeBench v6 (90.3%) and GPQA Diamond (89.2%), and does well on computer-use/mobile-use agentic benchmarks (OSWorld-Verified 84.3%, AndroidWorld 81.9%). It trails flagship models on frontier-exam reasoning (HLE 30.8% vs Qwen3.8 Max's 43.6%) and long-horizon professional-task simulation (JobBench 33.4%).

Specifications

FieldValue
OrganizationAlibaba
Parameters27B dense (causal LM with native vision encoder)
ArchitectureQwen 3.5 foundation
LicenseApache 2.0
Release dateAugust 2026
ModalityMultimodal (text, image, video)
Context window262,144 tokens native; extensible to 1,000,000 via YaRN

Pricing

Open weights available now on Hugging Face (Qwen/Qwen3.8-27B) under Apache 2.0 — free to self-host. A hosted version is planned via Qwen Cloud with 1M-token context and built-in tools by default ("coming soon" as of the model card).

Public Benchmark Scores

BenchmarkScoreSourceDate
Terminal Bench 2.173.0%Qwen3.8-27B model card2026-08
SWE-bench Pro61.7%Qwen3.8-27B model card2026-08
DeepSWE 1.142.2%Qwen3.8-27B model card2026-08
LiveCodeBench v60.903Qwen3.8-27B model card2026-08
GPQA Diamond89.2%Qwen3.8-27B model card2026-08
Humanity's Last Exam30.8%Qwen3.8-27B model card2026-08
IFBench79.5%Qwen3.8-27B model card2026-08
JobBench33.4%Qwen3.8-27B model card2026-08
Agents' Last Exam42.9 (score)Qwen3.8-27B model card2026-08
WildClawBench48.02Qwen3.8-27B model card2026-08
OSWorld-Verified84.3%Qwen3.8-27B model card2026-08
MathVision90.0%Qwen3.8-27B model card2026-08
BabyVision65.7%Qwen3.8-27B model card2026-08
CharXiv Reasoning (no tools)83.7%Qwen3.8-27B model card2026-08
CharXiv Reasoning90.2%Qwen3.8-27B model card2026-08
OmniDocBench 1.591.1%Qwen3.8-27B model card2026-08
RealWorldQA85.9%Qwen3.8-27B model card2026-08

Scores are self-reported by Alibaba's Qwen Team on the official Hugging Face model card (Qwen/Qwen3.8-27B), cross-checked against the card's structured HF "Eval Results" metadata where available (SWE-bench Pro, DeepSWE, GPQA Diamond, HLE, WildClawBench, ClawEval-MM all independently confirmed via that mechanism).

Notable Results Not Yet Tracked as Benchgen Benchmarks

Several benchmarks reported on the model card don't yet have standalone Benchgen pages — either because they're explicitly in-house/proprietary evaluations, or because Benchgen hasn't verified independent public documentation for them yet:

  • NL2Repo-Bench (repo-level code generation, 42.3%), WebArena-Verified (browser use, 64.8%), AndroidWorld (mobile use, 81.9%), ClawEval-MM (multimodal tool use, Pass@3 57.4%), and ERQA (embodied reasoning, 65.5%) appear to be real, independently-documented benchmarks but don't have Benchgen pages yet — flagged as a follow-up.
  • CoWorkBench (70.7%), QwenSWEBench (79.0%), and RecreationBench (47.1%) are explicitly labeled in-house/proprietary Qwen evaluations on the model card — not independently reproducible, so not listed as Benchgen leaderboard entries.
  • SWE-MM (38.6%) and Vision2Web (62.9%) are modified/non-standard variants of existing benchmark families (SWE-bench Multimodal, web-agent benchmarks respectively) without clear independent public documentation of the exact modifications — not added pending further verification.

Qwen3.8-27B vs Alternatives

ModelTerminal Bench 2.1GPQA DiamondHLELicense
Qwen3.8-27B73.0%89.2%30.8%Apache 2.0
Qwen3.8 Max86.6%92.6%43.6%Qwen3.8-Max License (custom)
Qwen3-8BApache 2.0

Qwen3.8-27B trails its own flagship sibling Qwen3.8 Max on every compared metric, as expected given the ~90x parameter gap (27B dense vs 2.4T/95B-active MoE) — but it offers native vision-language understanding and agentic computer-use capability (OSWorld-Verified 84.3%, AndroidWorld 81.9%) in a self-hostable, consumer-hardware-friendly package.

FAQ

Is Qwen3.8-27B open source? Yes — Apache 2.0, with weights available on Hugging Face (Qwen/Qwen3.8-27B).

How big is Qwen3.8-27B? 27 billion parameters, dense (not mixture-of-experts), with a native vision encoder for image and video understanding.

What is Qwen3.8-27B's context window? 262,144 tokens natively, extensible to 1,000,000 tokens via YaRN RoPE scaling.

Does Qwen3.8-27B support images and video? Yes — it's a native vision-language model, not a text-only model with a bolted-on adapter. It handles images (including STEM diagrams and documents) and hour-scale video understanding.

How does Qwen3.8-27B compare to Qwen3.8 Max? Qwen3.8-27B is the compact, dense sibling to the 2.4T/95B-active MoE flagship Qwen3.8 Max — it trails Qwen3.8 Max on every benchmark (e.g. 73.0% vs 86.6% on Terminal Bench 2.1) but is far smaller and fully open-weight under plain Apache 2.0.

Where can I access Qwen3.8-27B? Self-hosted from the open weights at huggingface.co/Qwen/Qwen3.8-27B, via SGLang/vLLM/TokenSpeed. A hosted Qwen Cloud version is planned ("coming soon" per the model card).


Benchmark scores sourced from the official Hugging Face model card for Qwen/Qwen3.8-27B (Aug 2026), cross-checked against the card's structured HF Eval Results metadata.