| Rank | Model | Score |
|---|---|---|
| 1 | gpt-5-2-pro-2025-12-11 | 0.898 |
| 2 | gpt-5 | 0.888 |
1 phaseActive
BrowseComp variant with 256K context window for evaluating agents finding hard-to-locate web information.
Quick answer: BrowseComp Long Context 256k evaluates web browsing agents on 1,266 questions requiring persistent internet navigation. Uses 256K context window, testing capacity for large-scale retrieval and reasoning.
| Property | Value |
|---|---|
| Tasks | 1,266 questions |
| Context | 256K tokens |
| Metric | Answer accuracy (short verifiable answers) |
| Parent benchmark | BrowseComp |
Source: Wei et al. 2025. Last updated 2026-07-24.