Benchgen

SuperGPQA — Results

RankModelScore
1deepseek-v4-1-flash53.1
2minicpm5-2b40.8
3qwen3-7-max0.736
4qwen3-6-plus0.716
5qwen3-7-plus0.714
6seed-2-1-pro0.708
7qwen3-5-397b-a17b0.704
8seed-2-1-turbo0.674
9qwen3-5-122b-a10b0.671
10qwen3-6-27b0.66
11qwen3-5-27b0.656
12qwen3-235b-a22b-thinking-25070.649
13qwen3-6-35b-a3b0.647
14qwen3-vl-235b-a22b-thinking0.643
15qwen3-5-35b-a3b0.634
16qwen3-235b-a22b-instruct-25070.626
17qwen3-next-80b-a3b-thinking0.608
18qwen3-vl-235b-a22b-instruct0.604
19qwen3-vl-32b-thinking0.59
20qwen3-next-80b-a3b-instruct0.588
21qwen3-5-9b0.582
22kimi-k2-instruct-09050.572
23kimi-k2-instruct0.572
24qwen3-vl-30b-a3b-thinking0.564
25qwen3-vl-32b-instruct0.546
S

SuperGPQA

1 phaseActive

25,957 graduate-level questions across 285 academic disciplines including Engineering, Medicine, Law, and Science. Expert-validated with 80+ annotators. Metric: accuracy.

Overview

SuperGPQA

Category Metric Tasks Saturation

Paper GitHub Dataset

Quick answer: SuperGPQA is a comprehensive graduate-level reasoning benchmark spanning 285 academic disciplines with 25,957 expert-validated questions. It extends the GPQA framework dramatically — from 448 questions in 3 domains to nearly 26,000 questions across Engineering, Medicine, Science, Law, and specialized fields. Questions are collaboratively filtered by over 80 expert annotators, making SuperGPQA one of the most thorough expert-knowledge evaluations available. With only 34 models evaluated as of mid-2026, it remains relatively uncrowded.

At a Glance

What it tests: Graduate-level academic knowledge and reasoning across a breadth of disciplines — from quantum mechanics to legal theory, from pathology to agricultural science. Questions require the kind of specialized domain knowledge that only subject matter experts reliably command.

Why it matters: GPQA Diamond (198 questions, 3 domains) is the current standard for expert-knowledge evaluation, but its small size creates high variance. SuperGPQA's 285 disciplines and 25,957 questions provide dramatically better statistical reliability and far broader disciplinary coverage.

Known limitations: Newer benchmark with fewer evaluation results than GPQA. Some niche disciplines may have limited public training data, making it hard to isolate genuine reasoning from memorization.

What SuperGPQA Measures

SuperGPQA employs a Human-LLM collaborative filtering mechanism: over 80 expert annotators from diverse academic backgrounds create and validate questions across 13 broad disciplinary areas:

  • Engineering (electrical, mechanical, civil, computer, chemical)
  • Medicine and Health Sciences
  • Natural Sciences (physics, chemistry, biology, geology)
  • Social Sciences (economics, psychology, sociology)
  • Humanities (philosophy, history, linguistics)
  • Law and Legal Studies
  • Light Industry, Agriculture, and Service-Oriented Domains

Questions are multiple-choice with 4 options. Each question is designed to require genuine domain expertise — not general reasoning or trivia knowledge.

Benchmark Specifications

FieldValue
Total questions25,957
Academic disciplines285
Broad areas13
Primary metricAccuracy (% correct)
Question formatMultiple choice (4 options)
Expert annotators80+
PaperarXiv:2502.14739 (Feb 2025)
SaturationLow — significant room for improvement

Frequently Asked Questions

How is SuperGPQA different from GPQA Diamond? GPQA Diamond has 198 questions across 3 domains (biology, chemistry, physics). SuperGPQA has 25,957 questions across 285 disciplines including engineering, law, medicine, agriculture, and many more. SuperGPQA provides much broader coverage and better statistical reliability.

What score do frontier models achieve on SuperGPQA? As of 2025–2026, top models score in the 50–70% range, compared to 85%+ on GPQA Diamond. The broader disciplinary coverage and niche specializations make SuperGPQA significantly harder than GPQA Diamond for most models.

Does SuperGPQA have a Chinese focus? SuperGPQA includes Chinese-language academic disciplines and was developed with international annotators. It is designed to be multilingual in scope, though English is the primary evaluation language.