| Rank | Model | Score |
|---|---|---|
| 1 | qwen3-5-397b-a17b | 0.703 |
| 2 | qwen3-5-122b-a10b | 0.699 |
| 3 | qwen3-5-35b-a3b | 0.695 |
| 4 | longcat-flash-thinking | 0.69 |
| 5 | glm-4-7 | 0.666 |
| 6 | deepseek-v3-2-thinking | 0.65 |
| 7 | deepseek-v3-2 | 0.65 |
| 8 | kimi-k2-thinking-0905 | 0.623 |
| 9 | qwen3-5-27b | 0.621 |
| 10 | deepseek-v3-1 | 0.492 |
| 11 | minimax-m2 | 0.485 |
| 12 | deepseek-v3-2-exp | 0.479 |
| 13 | deepseek-r1-0528 | 0.357 |
1 phaseActive
Chinese web browsing benchmark with 289 multi-hop questions across 11 domains, addressing Chinese web linguistic and infrastructural complexities.
Quick answer: BrowseComp-zh (Zhou et al., 2025) is a high-difficulty Chinese web browsing benchmark with 289 multi-hop questions spanning 11 domains (Film & TV, Technology, Medicine, History). Questions are reverse-engineered from short, verifiable answers requiring sophisticated reasoning on the Chinese web.
| Property | Value |
|---|---|
| Tasks | 289 multi-hop questions |
| Domains | 11 (Film & TV, Tech, Medicine, History, etc.) |
| Language | Chinese |
| Metric | Accuracy |
| Parent benchmark | BrowseComp |
Source: Zhou et al. 2025. Last updated 2026-07-24.