Benchgen

BrowseComp-zh — Results

RankModelScore
1minicpm5-2b43.5
2qwen3-5-397b-a17b0.703
3qwen3-5-122b-a10b0.699
4qwen3-5-35b-a3b0.695
5longcat-flash-thinking0.69
6glm-4-70.666
7deepseek-v3-2-thinking0.65
8deepseek-v3-20.65
9kimi-k2-thinking-09050.623
10qwen3-5-27b0.621
11deepseek-v3-10.492
12minimax-m20.485
13deepseek-v3-2-exp0.479
14deepseek-r1-05280.357
B

BrowseComp-zh

1 phaseActive

Chinese web browsing benchmark with 289 multi-hop questions across 11 domains, addressing Chinese web linguistic and infrastructural complexities.

Overview

BrowseComp-zh

Category Metric Tasks Saturation Language

Paper

Quick answer: BrowseComp-zh (Zhou et al., 2025) is a high-difficulty Chinese web browsing benchmark with 289 multi-hop questions spanning 11 domains (Film & TV, Technology, Medicine, History). Questions are reverse-engineered from short, verifiable answers requiring sophisticated reasoning on the Chinese web.

Benchmark Details

PropertyValue
Tasks289 multi-hop questions
Domains11 (Film & TV, Tech, Medicine, History, etc.)
LanguageChinese
MetricAccuracy
Parent benchmarkBrowseComp

Source: Zhou et al. 2025. Last updated 2026-07-24.