Benchgen

BrowseComp Long Context 128k — Results

RankModelScore
1gpt-5-2-pro-2025-12-110.92
2gpt-5-1-instant0.9
3gpt-5-1-thinking0.9
4gpt-5-10.9
5gpt-50.9
B

BrowseComp Long Context 128k

1 phaseActive

BrowseComp variant evaluating agents' ability to browse the web and find hard-to-locate information, with 128K long-context retrieval.

Overview

BrowseComp Long Context 128k

Category Metric Tasks Saturation

Paper

Quick answer: BrowseComp Long Context 128k evaluates web browsing agents on 1,266 questions requiring persistent internet navigation to find hard-to-locate, entangled information. Uses 128K context window for retrieval.

Benchmark Details

PropertyValue
Tasks1,266 questions
Context128K tokens
MetricAnswer accuracy (short verifiable answers)
DifficultyHard — obscure, time-invariant information
Parent benchmarkBrowseComp

Source: Wei et al. 2025. Last updated 2026-07-24.