1 phaseActive
The biology-subdomain subset of GPQA's graduate-level, Google-proof multiple-choice benchmark, covering molecular biology and genetics.
Quick answer: GPQA Biology is the biology-subdomain slice of GPQA, a dataset of graduate-level, "Google-proof" multiple-choice questions written and validated by PhD-level domain experts, isolating molecular biology and genetics questions so a model's biology reasoning can be assessed independently of its physics or chemistry performance.
What it tests: Whether a model can answer PhD-level molecular biology and genetics questions that are difficult even for highly skilled non-expert biologists (and physicists/chemists) with unrestricted internet access.
Why it matters: Aggregate GPQA scores can mask large per-domain gaps; isolating biology lets researchers see whether a model's overall GPQA score is being propped up by strength in physics or chemistry while biology reasoning lags — notably, biology was GPT-4's strongest domain in the original paper.
Known limitations: At 105 questions (from GPQA's 546-question extended set), the biology subset is GPQA's smallest domain slice, limiting statistical power for fine-grained model comparisons.
GPQA's full dataset spans biology, physics, and chemistry, with each question written by a domain expert and validated by other PhD-level experts to confirm it is both objective and "Google-proof" — meaning skilled non-experts with full internet access still cannot reliably answer it. The paper's domain breakdown (Table 3 of Rein et al. 2023) reports 105 biology questions in the 546-question extended set, split across two subdomains: Molecular Biology (85) and Genetics (20).
Because GPQA's biology questions skew heavily toward molecular biology, this subset is especially useful for teams that want to probe a model's understanding of cellular mechanisms, gene regulation, and molecular pathways specifically, rather than its aggregate science knowledge.
| Field | Value |
|---|---|
| Task category | Reasoning / graduate-level biology Q&A |
| Metric | Accuracy (%) on 4-choice multiple-choice questions |
| Number of tasks | 105 biology questions (extended-set breakdown) |
| Saturation | Medium — GPT-4 few-shot CoT scored 58.1% on biology questions in the original paper, its strongest of the three domains, versus PhD-level expert accuracy of 65%+ |
| Created by | David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, Samuel R. Bowman, et al. (NYU, Anthropic, Cohere) |
| Source paper | Rein et al. 2023 |
| GitHub | idavidrein/gpqa |
Each question is a 4-choice multiple-choice item; accuracy is the percentage of biology questions answered correctly. Random guessing yields 25% accuracy, PhD-level expert biologists average roughly 65% (74% when discounting acknowledged mistakes), and skilled non-expert validators with unrestricted internet access average only around 34%.
In the original paper, GPT-4 with few-shot chain-of-thought prompting scored 58.1% accuracy on biology questions — notably its best-performing domain, well above its physics (37.0%) and chemistry (31.8%) scores. Newer frontier reasoning models have since pushed GPQA Diamond accuracy well above 90%, but biology-specific breakdowns for the newest models are less commonly reported.
No Benchgen results yet — be the first to run GPQA Biology.
| Benchmark | What it tests | Tasks | Saturation |
|---|---|---|---|
| GPQA Biology | PhD-level biology reasoning, Google-proof | 105 | medium |
| GPQA Chemistry | PhD-level chemistry reasoning, Google-proof | 214 | medium |
| GPQA Diamond | Highest-quality cross-domain GPQA subset | 198 | high |
GPQA Biology differs from GPQA Diamond by isolating a single scientific domain rather than mixing physics, biology, and chemistry, making it better suited for diagnosing domain-specific weaknesses rather than reporting one aggregate science-reasoning score.
Benchgen lets teams evaluate their own model against the biology-specific slice of GPQA, surfacing whether strong aggregate GPQA performance is being driven by physics/chemistry strength while biology reasoning lags behind.