Benchgen

SuperGPQA — Results

RankModelScore
1qwen3-7-max0.736
2qwen3-6-plus0.716
3qwen3-7-plus0.714
4seed-2-1-pro0.708
5qwen3-5-397b-a17b0.704
6seed-2-1-turbo0.674
7qwen3-5-122b-a10b0.671
8qwen3-6-27b0.66
9qwen3-5-27b0.656
10qwen3-235b-a22b-thinking-25070.649
11qwen3-6-35b-a3b0.647
12qwen3-vl-235b-a22b-thinking0.643
13qwen3-5-35b-a3b0.634
14qwen3-235b-a22b-instruct-25070.626
15qwen3-next-80b-a3b-thinking0.608
16qwen3-vl-235b-a22b-instruct0.604
17qwen3-vl-32b-thinking0.59
18qwen3-next-80b-a3b-instruct0.588
19qwen3-5-9b0.582
20kimi-k2-instruct-09050.572
21kimi-k2-instruct0.572
22qwen3-vl-30b-a3b-thinking0.564
23qwen3-vl-32b-instruct0.546
24qwen3-vl-30b-a3b-instruct0.531
25qwen3-5-4b0.529
S

SuperGPQA

1 phaseActive

25,957 graduate-level questions across 285 academic disciplines including Engineering, Medicine, Law, and Science. Expert-validated with 80+ annotators. Metric: accuracy.

Overview

SuperGPQA

Category Metric Tasks Saturation

Paper GitHub Dataset

Quick answer: SuperGPQA is a comprehensive graduate-level reasoning benchmark spanning 285 academic disciplines with 25,957 expert-validated questions. It extends the GPQA framework dramatically — from 448 questions in 3 domains to nearly 26,000 questions across Engineering, Medicine, Science, Law, and specialized fields. Questions are collaboratively filtered by over 80 expert annotators, making SuperGPQA one of the most thorough expert-knowledge evaluations available. With only 34 models evaluated as of mid-2026, it remains relatively uncrowded.

At a Glance

What it tests: Graduate-level academic knowledge and reasoning across a breadth of disciplines — from quantum mechanics to legal theory, from pathology to agricultural science. Questions require the kind of specialized domain knowledge that only subject matter experts reliably command.

Why it matters: GPQA Diamond (198 questions, 3 domains) is the current standard for expert-knowledge evaluation, but its small size creates high variance. SuperGPQA's 285 disciplines and 25,957 questions provide dramatically better statistical reliability and far broader disciplinary coverage.

Known limitations: Newer benchmark with fewer evaluation results than GPQA. Some niche disciplines may have limited public training data, making it hard to isolate genuine reasoning from memorization.

What SuperGPQA Measures

SuperGPQA employs a Human-LLM collaborative filtering mechanism: over 80 expert annotators from diverse academic backgrounds create and validate questions across 13 broad disciplinary areas:

  • Engineering (electrical, mechanical, civil, computer, chemical)
  • Medicine and Health Sciences
  • Natural Sciences (physics, chemistry, biology, geology)
  • Social Sciences (economics, psychology, sociology)
  • Humanities (philosophy, history, linguistics)
  • Law and Legal Studies
  • Light Industry, Agriculture, and Service-Oriented Domains

Questions are multiple-choice with 4 options. Each question is designed to require genuine domain expertise — not general reasoning or trivia knowledge.

Benchmark Specifications

FieldValue
Total questions25,957
Academic disciplines285
Broad areas13
Primary metricAccuracy (% correct)
Question formatMultiple choice (4 options)
Expert annotators80+
PaperarXiv:2502.14739 (Feb 2025)
SaturationLow — significant room for improvement

Frequently Asked Questions

How is SuperGPQA different from GPQA Diamond? GPQA Diamond has 198 questions across 3 domains (biology, chemistry, physics). SuperGPQA has 25,957 questions across 285 disciplines including engineering, law, medicine, agriculture, and many more. SuperGPQA provides much broader coverage and better statistical reliability.

What score do frontier models achieve on SuperGPQA? As of 2025–2026, top models score in the 50–70% range, compared to 85%+ on GPQA Diamond. The broader disciplinary coverage and niche specializations make SuperGPQA significantly harder than GPQA Diamond for most models.

Does SuperGPQA have a Chinese focus? SuperGPQA includes Chinese-language academic disciplines and was developed with international annotators. It is designed to be multilingual in scope, though English is the primary evaluation language.