| Rank | Model | Score |
|---|---|---|
| 1 | qwen3-7-max | 0.736 |
| 2 | qwen3-6-plus | 0.716 |
| 3 | qwen3-7-plus | 0.714 |
| 4 | seed-2-1-pro | 0.708 |
| 5 | qwen3-5-397b-a17b | 0.704 |
| 6 | seed-2-1-turbo | 0.674 |
| 7 | qwen3-5-122b-a10b | 0.671 |
| 8 | qwen3-6-27b | 0.66 |
| 9 | qwen3-5-27b | 0.656 |
| 10 | qwen3-235b-a22b-thinking-2507 | 0.649 |
| 11 | qwen3-6-35b-a3b | 0.647 |
| 12 | qwen3-vl-235b-a22b-thinking | 0.643 |
| 13 | qwen3-5-35b-a3b | 0.634 |
| 14 | qwen3-235b-a22b-instruct-2507 | 0.626 |
| 15 | qwen3-next-80b-a3b-thinking | 0.608 |
| 16 | qwen3-vl-235b-a22b-instruct | 0.604 |
| 17 | qwen3-vl-32b-thinking | 0.59 |
| 18 | qwen3-next-80b-a3b-instruct | 0.588 |
| 19 | qwen3-5-9b | 0.582 |
| 20 | kimi-k2-instruct-0905 | 0.572 |
| 21 | kimi-k2-instruct | 0.572 |
| 22 | qwen3-vl-30b-a3b-thinking | 0.564 |
| 23 | qwen3-vl-32b-instruct | 0.546 |
| 24 | qwen3-vl-30b-a3b-instruct | 0.531 |
| 25 | qwen3-5-4b | 0.529 |
1 phaseActive
25,957 graduate-level questions across 285 academic disciplines including Engineering, Medicine, Law, and Science. Expert-validated with 80+ annotators. Metric: accuracy.
Quick answer: SuperGPQA is a comprehensive graduate-level reasoning benchmark spanning 285 academic disciplines with 25,957 expert-validated questions. It extends the GPQA framework dramatically — from 448 questions in 3 domains to nearly 26,000 questions across Engineering, Medicine, Science, Law, and specialized fields. Questions are collaboratively filtered by over 80 expert annotators, making SuperGPQA one of the most thorough expert-knowledge evaluations available. With only 34 models evaluated as of mid-2026, it remains relatively uncrowded.
What it tests: Graduate-level academic knowledge and reasoning across a breadth of disciplines — from quantum mechanics to legal theory, from pathology to agricultural science. Questions require the kind of specialized domain knowledge that only subject matter experts reliably command.
Why it matters: GPQA Diamond (198 questions, 3 domains) is the current standard for expert-knowledge evaluation, but its small size creates high variance. SuperGPQA's 285 disciplines and 25,957 questions provide dramatically better statistical reliability and far broader disciplinary coverage.
Known limitations: Newer benchmark with fewer evaluation results than GPQA. Some niche disciplines may have limited public training data, making it hard to isolate genuine reasoning from memorization.
SuperGPQA employs a Human-LLM collaborative filtering mechanism: over 80 expert annotators from diverse academic backgrounds create and validate questions across 13 broad disciplinary areas:
Questions are multiple-choice with 4 options. Each question is designed to require genuine domain expertise — not general reasoning or trivia knowledge.
| Field | Value |
|---|---|
| Total questions | 25,957 |
| Academic disciplines | 285 |
| Broad areas | 13 |
| Primary metric | Accuracy (% correct) |
| Question format | Multiple choice (4 options) |
| Expert annotators | 80+ |
| Paper | arXiv:2502.14739 (Feb 2025) |
| Saturation | Low — significant room for improvement |
How is SuperGPQA different from GPQA Diamond? GPQA Diamond has 198 questions across 3 domains (biology, chemistry, physics). SuperGPQA has 25,957 questions across 285 disciplines including engineering, law, medicine, agriculture, and many more. SuperGPQA provides much broader coverage and better statistical reliability.
What score do frontier models achieve on SuperGPQA? As of 2025–2026, top models score in the 50–70% range, compared to 85%+ on GPQA Diamond. The broader disciplinary coverage and niche specializations make SuperGPQA significantly harder than GPQA Diamond for most models.
Does SuperGPQA have a Chinese focus? SuperGPQA includes Chinese-language academic disciplines and was developed with international annotators. It is designed to be multilingual in scope, though English is the primary evaluation language.