Benchgen

GDPVal-AA v2 — Results

RankModelScore
1claude-fable-51760
2grok-4-61753
3gpt-5-6-sol1748
4kimi-k31686
5muse-spark-1-21631
6glm-5-21514
7deepseek-v4-pro1307
8solar-pro-41276
9inkling1238
10kimi-k2-61190
11nemotron-3-ultra-550b-a55b1164
12kimi-k2-51009
13gemini-3-1-pro962
14muse-glimmer953
15nemotron-3-5-lightning-30b-a3b865

GDPVal-AA v2

1 phaseActive

Artificial Analysis's Elo-rated benchmark for real-world work tasks across occupations. Human baseline = 1,000. Higher is better. Part of the AA Intelligence Index v4.1.

Overview

GDPVal-AA v2

Category Metric Baseline Saturation

Leaderboard

Quick answer: GDPVal-AA v2 is Artificial Analysis's benchmark for measuring AI model performance on real-world, economically valuable tasks across a wide range of occupations. Scores are Elo ratings anchored to a human baseline of 1,000 — higher scores indicate performance above the human baseline. It is one of the nine evaluations in the Artificial Analysis Intelligence Index v4.1. Claude Fable 5 leads at 1,760 Elo.

At a Glance

What it tests: An AI model's ability to complete real-world work tasks that have direct economic value — the kind of tasks performed by knowledge workers, professionals, and specialists across occupations. Tasks are drawn from realistic work scenarios rather than narrow academic exercises.

Why it matters: Most intelligence benchmarks test reasoning or knowledge in isolation. GDPVal-AA v2 specifically targets economically productive output — aligning AI evaluation with real-world utility. By anchoring to a human baseline of 1,000, the scores give intuitive meaning: a score of 1,300 means substantially better than a typical human at these tasks, while 900 means somewhat below.

Known limitations: GDPVal-AA v2 is a proprietary benchmark developed and maintained by Artificial Analysis. Methodology details are published at artificialanalysis.ai/methodology/intelligence-benchmarking. The benchmark tasks and exact scoring methodology are not open-source.

What GDPVal-AA v2 Measures

GDPVal-AA v2 evaluates AI models on real-world, economically valuable tasks across a wide range of occupations. This includes tasks like:

  • Writing, summarizing, and editing professional documents
  • Analyzing data and producing actionable recommendations
  • Planning, scheduling, and workflow management
  • Research and information synthesis
  • Professional communication and correspondence

Each task is scored on whether the model's output would be considered useful and correct by a domain expert, generating Elo ratings through comparative evaluation. The score is anchored so that the human baseline = 1,000.

The Intelligence Index uses a normalized form of the score: (Elo - 500) / 2000, but the raw Elo values reported by model cards are the standard comparison metric.

Benchmark Specifications

FieldValue
Task categoryAgentic / real-world work
MetricElo rating (anchored to human baseline = 1,000)
Score rangeTypically ~400–1,800 (open-ended Elo)
Human baseline1,000
Created byArtificial Analysis
Versionv2
Leaderboardartificialanalysis.ai/evaluations/gdpval-aa
Methodologyartificialanalysis.ai/methodology/intelligence-benchmarking

How GDPVal-AA v2 Is Scored

Models are evaluated on a set of real-world work tasks. Outputs are compared against a human baseline and against each other using Elo-style comparison, producing a rating anchored so that an average human worker scores approximately 1,000. Scores above 1,000 indicate superhuman performance on the task distribution; scores below 1,000 indicate below-human performance.

State-of-the-Art Results

Scores from Inkling model card (Thinking Machines Lab, July 2026). All 9 comparison models reported.

RankModelScore (Elo)vs. HumanWeights
1Claude Fable 51,760+76%Closed
2GPT-5.6 Sol1,748+75%Closed
3GLM 5.21,514+51%Open
4DeepSeek V4 Pro1,307+31%Open
5Inkling1,238+24%Open
6Kimi K2.61,190+19%Open
7Nemotron 3 Ultra 550B1,164+16%Open
8Kimi K2.51,009~0%Open
9Gemini 3.1 Pro962-4%Closed

Human baseline = 1,000.

BenchmarkFocusMetricReal tasks?
GDPVal-AA v2Real-world economic workElo vs. human baselineYes
τ2-Bench (Banking)Customer service agentTask success rateSimulated
MCP-AtlasMCP tool-use% pass rateReal MCP servers
TerminalBenchCLI tasks% pass rateSandboxed

Last updated 2026-07-16.