Benchgen

GPQA Diamond — Results

RankModelScore
1fugu95.5
2fugu-ultra95.5
3gpt-5-6-sol94.6
4gemini-3-1-pro94.3
5claude-opus-4-794.2
6claude-opus-4-893.6
7gpt-5-593.6
8kimi-k393.5
9gpt-5-2-pro-2025-12-1193.2
10grok-4-593
11gpt-5-6-terra92.9
12gpt-5-492.8
13qwen3-8-max92.6
14gpt-5-292.4
15qwen3-7-max92.4
16gpt-5-6-luna92.3
17grok-4-2092
18gemini-3-pro91.9
19claude-opus-4-691.3
20glm-5-291.2
21kimi-k2-690.5
22gemini-3-flash90.4
23hy390.4
24qwen3-6-plus90.4
25qwen3-7-plus90.3

GPQA Diamond

1 phaseActive

198-question graduate-level biology, chemistry, and physics benchmark — expert-validated, PhD-hard. Metric: % correct. Fugu/Fugu Ultra reach 95.5%.

Overview

GPQA Diamond

Category Metric Tasks Saturation Created

Paper GitHub Dataset

Quick answer: GPQA Diamond is a 198-question multiple-choice benchmark of PhD-level biology, chemistry, and physics questions, created by Rein et al. (2023). Questions are validated by domain experts and designed so that non-expert humans score below 40%; PhD-level domain experts score around 65–74%. As of June 2026, both Fugu and Fugu Ultra reach 95.5%, matching the frontier.

At a Glance

What it tests: Graduate-level scientific knowledge and reasoning across biology, chemistry, and physics — questions that require PhD-level domain expertise to answer correctly.

Why it matters: GPQA Diamond has become the industry standard for measuring graduate-level scientific reasoning. Every major lab reports it; passing it well (>85%) indicates genuine expert-level understanding, not pattern matching.

Known limitations: At 198 questions, GPQA Diamond has high score variance — a few questions answered differently can shift a model's score by 1–2 percentage points. Top models now cluster above 90%, and it is beginning to approach saturation at the frontier.

What GPQA Diamond Measures

GPQA (Graduate-Level Google-Proof Q&A) was designed to be difficult even for well-informed non-experts with internet access. The "Diamond" subset is the hardest tier — 198 questions where expert validators rated the question as having a clear, unambiguous correct answer that non-experts consistently fail on. Each question is four-choice multiple choice.

To ensure quality, each GPQA Diamond question was written by a PhD student or researcher and then validated by multiple domain experts who agreed it was (a) correct, (b) unambiguous, and (c) genuinely hard. Non-expert humans with PhDs in adjacent fields score around 34%; domain experts score 65–74%. This calibration makes GPQA Diamond one of the most informative hard-science benchmarks available.

By 2026, frontier models approach 95%+, suggesting the top of the leaderboard is beginning to saturate on this benchmark.

Benchmark Specifications

FieldValue
Task categoryReasoning / graduate-level science
Metric% correct (4-choice multiple choice)
Number of questions198 (Diamond subset)
Domain coverageBiology, chemistry, physics
Expert baseline65–74% (PhD domain experts)
Non-expert baseline~34%
SaturationMedium (frontier models now approach 95%)
Created byRein et al.
Source paperGPQA: A Graduate-Level Google-Proof Q&A Benchmark (2023)
GitHubidavidrein/gpqa
DatasetHuggingFace — Idavidrein/gpqa

How GPQA Diamond Is Scored

Each question is 4-choice multiple choice; a model scores 1 point per correct answer. The score is reported as a percentage. Due to the small 198-question set, standard error is relatively high (~3.5pp at 95% CI for a 50% score). Scores above 85% are considered strong; scores above 90% are frontier-class. Many labs use chain-of-thought prompting; some use majority voting or extended thinking, which can add a few percentage points.

State-of-the-Art Results

RankModelScoreSourceDate
1Fugu95.5%Sakana Fugu technical report2026-06
1Fugu Ultra95.5%Sakana Fugu technical report2026-06
3Fable 5 / Mythos Preview (max)92.0%Sakana Fugu technical report2026-06
4Gemini 3 Pro91.9%Google DeepMind2026-06
5Grok 487.5%xAI2025-07
6GPT-588.4%OpenAI2026-06

Scores attributed to respective model providers. Sakana Fugu scores use the evaluation conditions in the Fugu technical report.