Quick answer: Qwen3.8-Omni-Flash is a multimodal model in Alibaba's Qwen3.8 family, released via arXiv paper and QwenCloud API in September 2026. Unlike many Qwen releases, it is offered as a proprietary, API-only model rather than open weights — consistent with Alibaba's closed QwenCloud tier (alongside the similarly API-only Qwen3.8-Max). It scores 91.0% on GPQA Diamond, 92.6% on LiveCodeBench v6, and 80.5% on SWE-bench Multilingual.
Where it's strong: High coding and instruction-following scores — LiveCodeBench v6 (92.6%), IFBench (81.5%), SWE-bench Multilingual (80.5%).
Where it's weaker: Humanity's Last Exam (36.5%) trails frontier flagship models; NL2Repo-Bench (48.9%) suggests more limited repository-scale code generation than top coding specialists.
Licensing note: Alibaba's own paper/documentation describes some underlying components as open-source research artifacts, but Qwen3.8-Omni-Flash itself is distributed only via a closed QwenCloud API — treat the license as proprietary/API-only unless Alibaba publishes downloadable weights.
Qwen3.8-Omni-Flash is a multimodal model in Alibaba's Qwen3.8 generation, documented via an arXiv technical paper and made available through Alibaba's QwenCloud API rather than as downloadable open weights. This marks a departure from the open-weight releases Qwen is best known for — Qwen3.8-Omni-Flash instead sits alongside Qwen3.8-Max as a closed, API-only offering, reflecting Alibaba's dual-track strategy of open-weight models for the research community and proprietary, higher-margin API tiers for production use.
The model's benchmark profile emphasizes coding and structured instruction-following: strong results on LiveCodeBench v6, SWE-bench (both Pro and Multilingual variants), and IFBench, alongside solid general reasoning (GPQA Diamond 91.0%) but a more modest result on the extreme-difficulty Humanity's Last Exam (36.5%), consistent with a model tuned for practical coding/agentic workloads rather than maximal frontier reasoning.
| Field | Value |
|---|---|
| Organization | Alibaba (Qwen team) |
| License | Proprietary (API only — closed QwenCloud tier) |
| Modality | Multimodal |
| Release date | 2026-09 |
| Availability | QwenCloud API |
| Benchmark | Score | Source | Date |
|---|---|---|---|
| GPQA Diamond | 91.0% | Alibaba Qwen3.8-Omni-Flash technical paper (arXiv) | 2026-09 |
| Humanity's Last Exam | 36.5% | Alibaba Qwen3.8-Omni-Flash technical paper (arXiv) | 2026-09 |
| SWE-bench Pro | 63.3% | Alibaba Qwen3.8-Omni-Flash technical paper (arXiv) | 2026-09 |
| SWE-bench Multilingual | 80.5% | Alibaba Qwen3.8-Omni-Flash technical paper (arXiv) | 2026-09 |
| NL2Repo-Bench | 48.9% | Alibaba Qwen3.8-Omni-Flash technical paper (arXiv) | 2026-09 |
| IFBench | 81.5% | Alibaba Qwen3.8-Omni-Flash technical paper (arXiv) | 2026-09 |
| LiveCodeBench v6 | 92.6% | Alibaba Qwen3.8-Omni-Flash technical paper (arXiv) | 2026-09 |
Scores above are self-reported by Alibaba's Qwen3.8-Omni-Flash technical paper and shown for context; not independent Benchgen measurements. The paper also reports a "DeepSWE 1.1" result and a "CoWorkBench" result; the former is not added here pending confirmation it matches Benchgen's existing DeepSWE benchmark version, and the latter is omitted entirely as Benchgen has no corresponding benchmark page.