Quick answer: HY3 is Hungyuan's July 2026 model scoring 78.0% on SWE-Bench Verified and 28.8% on LHTB. Its SWE-Bench Verified score places it among the strongest models for real-world software engineering tasks.
Where HY3 leads
Where it lags
Best for: Real-world software engineering agent tasks; SWE-Bench class code-fixing pipelines.
HY3 is Hungyuan's July 2026 model. Its primary distinguishing feature is a 78.0% SWE-Bench Verified score — placing it at near-parity with DeepSeek-V4 Flash Max (79%) and above GPT-5 Codex (74.5%) on this real-world software engineering benchmark.
Limited public benchmark data is available for HY3 beyond SWE-Bench and LHTB. The 28.8% LHTB (Long Horizon Task Benchmark) suggests more moderate performance on complex multi-step agentic tasks compared to its SWE-Bench strength.
| Field | Value |
|---|---|
| Organization | Hungyuan |
| License | Proprietary (API only) |
| Release date | July 7, 2026 |
| Knowledge cutoff | April 2026 |
| Modality | Text |
Available via Hungyuan API. Refer to official Hungyuan pricing for current rates.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| SWE-Bench Verified | 78.0% | Benchgen evaluation | 2026-07 |
| LHTB | 28.8% | Benchgen evaluation | 2026-07 |
| Model | SWE-Bench Verified | License |
|---|---|---|
| HY3 | 78.0% | Proprietary |
| DeepSeek-V4 Flash Max | 79% | Proprietary |
| GPT-5 Codex | 74.5% | Proprietary |
| DeepSeek-V3.2 Speciale | 73.1% | MIT |
HY3 is competitive with DeepSeek-V4 Flash Max on SWE-Bench (78% vs 79%). For open-weight SWE alternatives: DeepSeek-V3.2 Speciale (73.1%, MIT).
Specs from Hungyuan's official HY3 release (July 2026) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.