Quick answer: Hy4 preview is Tencent Hunyuan's August 2026 flagship — a 770B-parameter Mixture-of-Experts model with 49B active parameters per token, a 1M-token context window, and Apache 2.0 open weights. It scores 92.3% on GPQA Diamond and 82.9% on SWE-bench Multilingual, a major generation-over-generation jump from Hy3.
Where Hy4 preview leads
Where it lags
Best for: Teams that need a frontier-class, fully open-weight (Apache 2.0) model for agentic coding, tool use, and long-context workloads without API lock-in.
Hy4 preview is the newest flagship in Tencent Hunyuan's "Hy" model line, succeeding Hy3 preview. It's a Mixture-of-Experts model with 78 backbone layers (the first dense, the remaining 77 MoE with 256 routed experts + 1 shared expert, top-8 routed experts active per token), plus a native multi-token-prediction (MTP) layer for speculative decoding. The architecture borrows ideas from DeepSeek and GLM: attention uses Gated DeepSeek Sparse Attention (Gated DSA) with IndexCache for cross-layer sparse index reuse, and the residual stream uses identity Hyper-Connections (iHC) to widen inter-layer information flow.
Tencent frames this release around agentic productivity rather than raw knowledge benchmarks alone — the team built training data directly from internal software engineers, game developers, finance analysts, and security experts, and reports the largest generation-over-generation gain they've measured, particularly in agentic coding (SWE-bench Multilingual jumped from Hy3's 75.8% to 82.9%; DeepSWE from 28.0% to 64.3%). In a blind internal side-by-side (163 experts rating 203 engineering tasks), Hy4 preview edged out both GLM 5.3 and Kimi K3.
This is an early "preview" release — Tencent explicitly flags known issues (over-long reasoning chains, excessive self-verification) and says they expect to iterate quickly, the same pattern that took Hy3 preview from launch to a substantially stronger model.
| Field | Value |
|---|---|
| Organization | Hungyuan (Tencent) |
| Architecture | Mixture-of-Experts (MoE), Gated DSA attention |
| Total parameters | 770B |
| Activated parameters | 49B |
| Layers | 78 (1 dense + 77 MoE) |
| Routed / shared experts | 256 / 1 (top-8 routed active per token) |
| Context length | 1M tokens |
| Vocabulary size | 120,832 |
| License | Apache 2.0 (open weights) |
| Release date | August 2026 |
| Modality | Text |
Open weights under Apache 2.0 — self-hosted via vLLM or SGLang (official prebuilt images vllm/vllm-openai:hy4-preview and lmsysorg/sglang:hy4-preview). No API list price; cost is your own inference infrastructure. An FP8-quantized variant (Hy4-preview-FP8) is also published for lower-memory deployment.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| GPQA Diamond | 92.3% | Tencent Hunyuan model card | 2026-08 |
| Humanity's Last Exam | 43.4% | Tencent Hunyuan model card | 2026-08 |
| SWE-bench Multilingual | 82.9% | Tencent Hunyuan model card | 2026-08 |
| SWE-bench Pro | 65.7% | Tencent Hunyuan model card | 2026-08 |
| DeepSWE | 64.3% | Tencent Hunyuan model card | 2026-08 |
| SWE-Marathon | 31.9% | Tencent Hunyuan model card | 2026-08 |
| Terminal-Bench 2.1 | 85.4% | Tencent Hunyuan model card | 2026-08 |
| NL2Repo-Bench | 58.9% | Tencent Hunyuan model card | 2026-08 |
| CyberGym | 78.4% | Tencent Hunyuan model card | 2026-08 |
| ProgramBench | 17.5% | Tencent Hunyuan model card | 2026-08 |
| WideSearch | 83.9% | Tencent Hunyuan model card | 2026-08 |
| OfficeQA Pro | 66.2% | Tencent Hunyuan model card | 2026-08 |
| MCP-Atlas | 83.7% | Tencent Hunyuan model card | 2026-08 |
| Toolathlon-Verified | 74.1% | Tencent Hunyuan model card | 2026-08 |
| APEX-Agents | 37.1% | Tencent Hunyuan model card | 2026-08 |
| JobBench | 61.7% | Tencent Hunyuan model card | 2026-08 |
| GDPval-AA V2 | 1678 (Elo) | Tencent Hunyuan model card | 2026-08 |
| CritPt | 16.9% | Tencent Hunyuan model card | 2026-08 |
| SWE Atlas — Codebase Q&A | 64.0% | Tencent Hunyuan model card | 2026-08 |
| SWE Atlas — Test Writing | 57.8% | Tencent Hunyuan model card | 2026-08 |
| SWE Atlas — Refactoring | 53.3% | Tencent Hunyuan model card | 2026-08 |
| Harbor-Index | 39.6% | Tencent Hunyuan model card | 2026-08 |
| BankerToolBench | 78.6% | Tencent Hunyuan model card | 2026-08 |
| SUPERChem | 66.4% | Tencent Hunyuan model card | 2026-08 |
| ArXivMath | 66.6% | Tencent Hunyuan model card | 2026-08 |
| HorizonMath | 8.8% | Tencent Hunyuan model card | 2026-08 |
| MathArena Apex 2025 | 74.2% | Tencent Hunyuan model card | 2026-08 |
| BrokenArXiv | 54.6% | Tencent Hunyuan model card | 2026-08 |
Scores as self-reported by Tencent Hunyuan for Hy4 preview's own column in the model card's benchmark appendix table (some third-party comparison scores in that same table are Tencent's own re-tests). Several additional benchmarks in the model card (AutomationBench v1.0.6, PostTrainBench V1.1, SkillsBench 79-task text-only subset, Agents' Last Exam ALE-CLI, HLE with-tools variant, and several Tencent-internal-only evals) are not listed here — see tasks/hy4-preview-progress.md for why each was excluded (version/subset mismatch with the existing tracked benchmark, or no independent public benchmark page).
| Model | GPQA Diamond | SWE-bench Multilingual | License |
|---|---|---|---|
| Hy4 Preview | 92.3% | 82.9% | Apache 2.0 |
| Hy3 | 90.9% | 75.8% | Proprietary |
| Kimi K3 | 93.5%/92.8%* | 80.8% | Apache 2.0 |
| GLM 5.3 | 91.7%/91.4%* | 81.3% | Proprietary (open weights ~2 weeks post-launch) |
| Qwen3.8 Max | 92.6%/92.2%* | 82.6% | Custom permissive commercial |
| DeepSeek V4 Pro | 92.8%/91.7%* | 77.3% | MIT |
| GPT 5.6 Sol | 94.1%/94.7%* | 74.1% | Proprietary |
| Claude Opus 5 | 93.7%/93.3%* | 89.5%/85.8%* | Proprietary |
Comparison scores marked * are Tencent's own re-tests of the other vendors' models (from the same model-card appendix table), not those vendors' own self-reported numbers — see tasks/hy4-preview-progress.md for the full sourcing note.
Hy4 preview leads its own generation (Hy3) by a wide margin on both agentic coding and knowledge benchmarks, and is competitive with — though not consistently ahead of — the strongest proprietary models (GPT 5.6 Sol, Claude Opus 5) while being the only model in this set with fully open Apache 2.0 weights at this parameter scale.
Specs from Tencent Hunyuan's official Hy4-preview Hugging Face model card and GitHub repository (August 2026). Last updated 2026-08-31.