Benchgen

GPQA Diamond — Results

RankModelScore
1gpt-6-astra96
2fugu95.5
3fugu-ultra95.5
4gpt-5-6-sol94.6
5gemini-3-1-pro94.3
6claude-opus-4-794.2
7claude-opus-4-893.6
8gpt-5-593.6
9kimi-k393.5
10gpt-5-2-pro-2025-12-1193.2
11grok-4-593
12gpt-5-6-terra92.9
13gpt-5-492.8
14ornith-1-5-397b92.8
15qwen3-8-max92.6
16gpt-5-292.4
17qwen3-7-max92.4
18gpt-5-6-luna92.3
19hy4-preview92.3
20grok-4-2092
21gemini-3-pro91.9
22qwen3-8-flash-next91.7
23claude-opus-4-691.3
24glm-5-291.2
25deepseek-v4-1-flash90.9

GPQA Diamond

1 phaseActive

198-question graduate-level biology, chemistry, and physics benchmark — expert-validated, PhD-hard. Metric: % correct. Fugu/Fugu Ultra reach 95.5%.

Overview

GPQA Diamond

Category Metric Tasks Saturation Created

Paper GitHub Dataset

Quick answer: GPQA Diamond is a 198-question multiple-choice benchmark of PhD-level biology, chemistry, and physics questions, created by Rein et al. (2023). Questions are validated by domain experts and designed so that non-expert humans score below 40%; PhD-level domain experts score around 65–74%. As of June 2026, both Fugu and Fugu Ultra reach 95.5%, matching the frontier.

At a Glance

What it tests: Graduate-level scientific knowledge and reasoning across biology, chemistry, and physics — questions that require PhD-level domain expertise to answer correctly.

Why it matters: GPQA Diamond has become the industry standard for measuring graduate-level scientific reasoning. Every major lab reports it; passing it well (>85%) indicates genuine expert-level understanding, not pattern matching.

Known limitations: At 198 questions, GPQA Diamond has high score variance — a few questions answered differently can shift a model's score by 1–2 percentage points. Top models now cluster above 90%, and it is beginning to approach saturation at the frontier.

What GPQA Diamond Measures

GPQA (Graduate-Level Google-Proof Q&A) was designed to be difficult even for well-informed non-experts with internet access. The "Diamond" subset is the hardest tier — 198 questions where expert validators rated the question as having a clear, unambiguous correct answer that non-experts consistently fail on. Each question is four-choice multiple choice.

To ensure quality, each GPQA Diamond question was written by a PhD student or researcher and then validated by multiple domain experts who agreed it was (a) correct, (b) unambiguous, and (c) genuinely hard. Non-expert humans with PhDs in adjacent fields score around 34%; domain experts score 65–74%. This calibration makes GPQA Diamond one of the most informative hard-science benchmarks available.

By 2026, frontier models approach 95%+, suggesting the top of the leaderboard is beginning to saturate on this benchmark.

Benchmark Specifications

FieldValue
Task categoryReasoning / graduate-level science
Metric% correct (4-choice multiple choice)
Number of questions198 (Diamond subset)
Domain coverageBiology, chemistry, physics
Expert baseline65–74% (PhD domain experts)
Non-expert baseline~34%
SaturationMedium (frontier models now approach 95%)
Created byRein et al.
Source paperGPQA: A Graduate-Level Google-Proof Q&A Benchmark (2023)
GitHubidavidrein/gpqa
DatasetHuggingFace — Idavidrein/gpqa

How GPQA Diamond Is Scored

Each question is 4-choice multiple choice; a model scores 1 point per correct answer. The score is reported as a percentage. Due to the small 198-question set, standard error is relatively high (~3.5pp at 95% CI for a 50% score). Scores above 85% are considered strong; scores above 90% are frontier-class. Many labs use chain-of-thought prompting; some use majority voting or extended thinking, which can add a few percentage points.

State-of-the-Art Results

RankModelScoreSourceDate
1Fugu95.5%Sakana Fugu technical report2026-06
1Fugu Ultra95.5%Sakana Fugu technical report2026-06
3Fable 5 / Mythos Preview (max)92.0%Sakana Fugu technical report2026-06
4Gemini 3 Pro91.9%Google DeepMind2026-06
5Grok 487.5%xAI2025-07
6GPT-588.4%OpenAI2026-06

Scores attributed to respective model providers. Sakana Fugu scores use the evaluation conditions in the Fugu technical report.