| Rank | Model | Score |
|---|---|---|
| 1 | gpt-5-2-pro-2025-12-11 | 0.92 |
| 2 | gpt-5-1-instant | 0.9 |
| 3 | gpt-5-1-thinking | 0.9 |
| 4 | gpt-5-1 | 0.9 |
| 5 | gpt-5 | 0.9 |
1 phaseActive
BrowseComp variant evaluating agents' ability to browse the web and find hard-to-locate information, with 128K long-context retrieval.
Quick answer: BrowseComp Long Context 128k evaluates web browsing agents on 1,266 questions requiring persistent internet navigation to find hard-to-locate, entangled information. Uses 128K context window for retrieval.
| Property | Value |
|---|---|
| Tasks | 1,266 questions |
| Context | 128K tokens |
| Metric | Answer accuracy (short verifiable answers) |
| Difficulty | Hard — obscure, time-invariant information |
| Parent benchmark | BrowseComp |
Source: Wei et al. 2025. Last updated 2026-07-24.