| Rank | Model | Score |
|---|---|---|
| 1 | claude-fable-5 | 1760 |
| 2 | grok-4-6 | 1753 |
| 3 | gpt-5-6-sol | 1748 |
| 4 | kimi-k3 | 1686 |
| 5 | muse-spark-1-2 | 1631 |
| 6 | glm-5-2 | 1514 |
| 7 | deepseek-v4-pro | 1307 |
| 8 | solar-pro-4 | 1276 |
| 9 | inkling | 1238 |
| 10 | kimi-k2-6 | 1190 |
| 11 | nemotron-3-ultra-550b-a55b | 1164 |
| 12 | kimi-k2-5 | 1009 |
| 13 | gemini-3-1-pro | 962 |
| 14 | muse-glimmer | 953 |
| 15 | nemotron-3-5-lightning-30b-a3b | 865 |
1 phaseActive
Artificial Analysis's Elo-rated benchmark for real-world work tasks across occupations. Human baseline = 1,000. Higher is better. Part of the AA Intelligence Index v4.1.
Quick answer: GDPVal-AA v2 is Artificial Analysis's benchmark for measuring AI model performance on real-world, economically valuable tasks across a wide range of occupations. Scores are Elo ratings anchored to a human baseline of 1,000 — higher scores indicate performance above the human baseline. It is one of the nine evaluations in the Artificial Analysis Intelligence Index v4.1. Claude Fable 5 leads at 1,760 Elo.
What it tests: An AI model's ability to complete real-world work tasks that have direct economic value — the kind of tasks performed by knowledge workers, professionals, and specialists across occupations. Tasks are drawn from realistic work scenarios rather than narrow academic exercises.
Why it matters: Most intelligence benchmarks test reasoning or knowledge in isolation. GDPVal-AA v2 specifically targets economically productive output — aligning AI evaluation with real-world utility. By anchoring to a human baseline of 1,000, the scores give intuitive meaning: a score of 1,300 means substantially better than a typical human at these tasks, while 900 means somewhat below.
Known limitations: GDPVal-AA v2 is a proprietary benchmark developed and maintained by Artificial Analysis. Methodology details are published at artificialanalysis.ai/methodology/intelligence-benchmarking. The benchmark tasks and exact scoring methodology are not open-source.
GDPVal-AA v2 evaluates AI models on real-world, economically valuable tasks across a wide range of occupations. This includes tasks like:
Each task is scored on whether the model's output would be considered useful and correct by a domain expert, generating Elo ratings through comparative evaluation. The score is anchored so that the human baseline = 1,000.
The Intelligence Index uses a normalized form of the score: (Elo - 500) / 2000, but the raw Elo values reported by model cards are the standard comparison metric.
| Field | Value |
|---|---|
| Task category | Agentic / real-world work |
| Metric | Elo rating (anchored to human baseline = 1,000) |
| Score range | Typically ~400–1,800 (open-ended Elo) |
| Human baseline | 1,000 |
| Created by | Artificial Analysis |
| Version | v2 |
| Leaderboard | artificialanalysis.ai/evaluations/gdpval-aa |
| Methodology | artificialanalysis.ai/methodology/intelligence-benchmarking |
Models are evaluated on a set of real-world work tasks. Outputs are compared against a human baseline and against each other using Elo-style comparison, producing a rating anchored so that an average human worker scores approximately 1,000. Scores above 1,000 indicate superhuman performance on the task distribution; scores below 1,000 indicate below-human performance.
Scores from Inkling model card (Thinking Machines Lab, July 2026). All 9 comparison models reported.
| Rank | Model | Score (Elo) | vs. Human | Weights |
|---|---|---|---|---|
| 1 | Claude Fable 5 | 1,760 | +76% | Closed |
| 2 | GPT-5.6 Sol | 1,748 | +75% | Closed |
| 3 | GLM 5.2 | 1,514 | +51% | Open |
| 4 | DeepSeek V4 Pro | 1,307 | +31% | Open |
| 5 | Inkling | 1,238 | +24% | Open |
| 6 | Kimi K2.6 | 1,190 | +19% | Open |
| 7 | Nemotron 3 Ultra 550B | 1,164 | +16% | Open |
| 8 | Kimi K2.5 | 1,009 | ~0% | Open |
| 9 | Gemini 3.1 Pro | 962 | -4% | Closed |
Human baseline = 1,000.
| Benchmark | Focus | Metric | Real tasks? |
|---|---|---|---|
| GDPVal-AA v2 | Real-world economic work | Elo vs. human baseline | Yes |
| τ2-Bench (Banking) | Customer service agent | Task success rate | Simulated |
| MCP-Atlas | MCP tool-use | % pass rate | Real MCP servers |
| TerminalBench | CLI tasks | % pass rate | Sandboxed |
Last updated 2026-07-16.