| Rank | Model | Score |
|---|---|---|
| 1 | fugu | 95.5 |
| 2 | fugu-ultra | 95.5 |
| 3 | gpt-5-6-sol | 94.6 |
| 4 | gemini-3-1-pro | 94.3 |
| 5 | claude-opus-4-7 | 94.2 |
| 6 | claude-opus-4-8 | 93.6 |
| 7 | gpt-5-5 | 93.6 |
| 8 | kimi-k3 | 93.5 |
| 9 | gpt-5-2-pro-2025-12-11 | 93.2 |
| 10 | grok-4-5 | 93 |
| 11 | gpt-5-6-terra | 92.9 |
| 12 | gpt-5-4 | 92.8 |
| 13 | qwen3-8-max | 92.6 |
| 14 | gpt-5-2 | 92.4 |
| 15 | qwen3-7-max | 92.4 |
| 16 | gpt-5-6-luna | 92.3 |
| 17 | grok-4-20 | 92 |
| 18 | gemini-3-pro | 91.9 |
| 19 | claude-opus-4-6 | 91.3 |
| 20 | glm-5-2 | 91.2 |
| 21 | kimi-k2-6 | 90.5 |
| 22 | gemini-3-flash | 90.4 |
| 23 | hy3 | 90.4 |
| 24 | qwen3-6-plus | 90.4 |
| 25 | qwen3-7-plus | 90.3 |
1 phaseActive
198-question graduate-level biology, chemistry, and physics benchmark — expert-validated, PhD-hard. Metric: % correct. Fugu/Fugu Ultra reach 95.5%.
Quick answer: GPQA Diamond is a 198-question multiple-choice benchmark of PhD-level biology, chemistry, and physics questions, created by Rein et al. (2023). Questions are validated by domain experts and designed so that non-expert humans score below 40%; PhD-level domain experts score around 65–74%. As of June 2026, both Fugu and Fugu Ultra reach 95.5%, matching the frontier.
What it tests: Graduate-level scientific knowledge and reasoning across biology, chemistry, and physics — questions that require PhD-level domain expertise to answer correctly.
Why it matters: GPQA Diamond has become the industry standard for measuring graduate-level scientific reasoning. Every major lab reports it; passing it well (>85%) indicates genuine expert-level understanding, not pattern matching.
Known limitations: At 198 questions, GPQA Diamond has high score variance — a few questions answered differently can shift a model's score by 1–2 percentage points. Top models now cluster above 90%, and it is beginning to approach saturation at the frontier.
GPQA (Graduate-Level Google-Proof Q&A) was designed to be difficult even for well-informed non-experts with internet access. The "Diamond" subset is the hardest tier — 198 questions where expert validators rated the question as having a clear, unambiguous correct answer that non-experts consistently fail on. Each question is four-choice multiple choice.
To ensure quality, each GPQA Diamond question was written by a PhD student or researcher and then validated by multiple domain experts who agreed it was (a) correct, (b) unambiguous, and (c) genuinely hard. Non-expert humans with PhDs in adjacent fields score around 34%; domain experts score 65–74%. This calibration makes GPQA Diamond one of the most informative hard-science benchmarks available.
By 2026, frontier models approach 95%+, suggesting the top of the leaderboard is beginning to saturate on this benchmark.
| Field | Value |
|---|---|
| Task category | Reasoning / graduate-level science |
| Metric | % correct (4-choice multiple choice) |
| Number of questions | 198 (Diamond subset) |
| Domain coverage | Biology, chemistry, physics |
| Expert baseline | 65–74% (PhD domain experts) |
| Non-expert baseline | ~34% |
| Saturation | Medium (frontier models now approach 95%) |
| Created by | Rein et al. |
| Source paper | GPQA: A Graduate-Level Google-Proof Q&A Benchmark (2023) |
| GitHub | idavidrein/gpqa |
| Dataset | HuggingFace — Idavidrein/gpqa |
Each question is 4-choice multiple choice; a model scores 1 point per correct answer. The score is reported as a percentage. Due to the small 198-question set, standard error is relatively high (~3.5pp at 95% CI for a 50% score). Scores above 85% are considered strong; scores above 90% are frontier-class. Many labs use chain-of-thought prompting; some use majority voting or extended thinking, which can add a few percentage points.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Fugu | 95.5% | Sakana Fugu technical report | 2026-06 |
| 1 | Fugu Ultra | 95.5% | Sakana Fugu technical report | 2026-06 |
| 3 | Fable 5 / Mythos Preview (max) | 92.0% | Sakana Fugu technical report | 2026-06 |
| 4 | Gemini 3 Pro | 91.9% | Google DeepMind | 2026-06 |
| 5 | Grok 4 | 87.5% | xAI | 2025-07 |
| 6 | GPT-5 | 88.4% | OpenAI | 2026-06 |
Scores attributed to respective model providers. Sakana Fugu scores use the evaluation conditions in the Fugu technical report.