Benchgen

Artificial Analysis Intelligence Index — Results

RankModelScore
1gpt-6-astra61.2
A

Artificial Analysis Intelligence Index

1 phaseActive

Artificial Analysis's composite index across agents, coding, general knowledge, and scientific reasoning — weighted average of 10 evaluations.

Overview

Artificial Analysis Intelligence Index

Category Metric Saturation Created

Methodology

Quick answer: The Artificial Analysis Intelligence Index is a composite score combining 10 underlying evaluations — spanning agentic tasks (30% weight), coding (20%), scientific reasoning (20%), and general knowledge (30%) — into a single number for comparing overall model capability. As of v4.3, the suite includes AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, AA-Omniscience, GDP.pdf, AA-LCR v1.1, Humanity's Last Exam, and CritPt.

At a Glance

What it tests: Overall model capability across four weighted categories — Agents (30%), Coding (20%), Scientific Reasoning (20%), General (30%) — synthesized from 10 individual evaluation datasets into one comparable number.

Why it matters: No single benchmark captures overall model quality; the Index is designed as the best available one-number synthesis for cross-model comparison, deliberately weighted toward agentic capability (30%) to reflect real-world deployment patterns.

Known limitations: As a composite, the Index can mask category-specific weaknesses — a model can score well overall while lagging badly in one category. It's also version-specific (methodology and constituent evaluations change between versions, e.g. v4.1.1 vs v4.3), so scores are only directly comparable within the same Index version.

What the Artificial Analysis Intelligence Index Measures

The Index combines a suite of evaluation datasets designed to assess language model capabilities across reasoning, knowledge, math, and programming into one weighted composite score. Artificial Analysis positions it as a more useful cross-model comparison than any single existing metric, while being explicit that — like all such syntheses — it has limitations and may not map directly onto every use case.

As of v4.3, the Index incorporates 10 evaluations across four categories: Agents (30% total weight: AA-Briefcase 15%, GDPval-AA v2 10%, AutomationBench-AA 5%), Coding (20%: Terminal-Bench v4.0 10%, SciCode 10%), General (30%: AA-Omniscience 15%, GDP.pdf 10%, AA-LCR v1.1 5%), and Scientific Reasoning (20%: Humanity's Last Exam 10%, CritPt 10%). Artificial Analysis reports a 95% confidence interval of under ±1% for the composite, based on repeated-run experiments, though individual constituent evaluations may have wider confidence intervals.

It's a primarily text-based, English-language evaluation suite — image inputs, speech inputs, and multilingual performance are benchmarked separately and are not part of the Index score itself.

Benchmark Specifications

FieldValue
Task categoryComposite / Reasoning
MetricWeighted-average composite index score
Category weightsAgents 30%, General 30%, Coding 20%, Scientific Reasoning 20%
Constituent evaluations (v4.3)AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, AA-Omniscience, GDP.pdf, AA-LCR v1.1, Humanity's Last Exam, CritPt
SaturationLow — frontier models still in the 50-65 range
Created byArtificial Analysis
Methodologyartificialanalysis.ai/methodology/intelligence-benchmarking

How the Artificial Analysis Intelligence Index Is Scored

Each constituent evaluation produces its own score (accuracy, pass@1, Elo-based pairwise comparison, etc., depending on the eval), which Artificial Analysis then combines into a single weighted-average composite using the category weights above. Scores are typically reported on a scale similar to a percentage (roughly 0-100), though the composite is best read as a relative ranking tool rather than an absolute accuracy figure, since it blends fundamentally different scoring types (accuracy, Elo, pass-rate) into one number.

State-of-the-Art Results

RankModelScoreVersionSourceDate
1Claude Fable 5.165.7v4.1.1OpenAI: GPT-6 Astra2026-09
2Claude Opus 563.1v4.1.1OpenAI: GPT-6 Astra2026-09
3GPT-6 Astra61.2v4.1.1OpenAI: GPT-6 Astra2026-09
4Claude Fable 562.1v4.1.1OpenAI: GPT-6 Astra2026-09
5GPT-5.6 Sol60.9v4.1.1OpenAI: GPT-6 Astra2026-09
6Gemini 3 Flash58.7v4.1.1OpenAI: GPT-6 Astra2026-09

Scores above are on Index v4.1.1 as reported in OpenAI's GPT-6 Astra announcement (September 2026) — note this predates the v4.3 methodology described in Artificial Analysis's own published methodology (10-evaluation set with AA-Briefcase/GDP.pdf added). Scores across different Index versions are not directly comparable.

Artificial Analysis Intelligence Index on Benchgen

No Benchgen results yet — be the first to run the Artificial Analysis Intelligence Index suite.

Artificial Analysis Intelligence Index vs Other Benchmarks

BenchmarkWhat it testsSaturation
AA Intelligence IndexComposite: agents, coding, general, scientific reasoningLow
AA Coding Agent IndexCoding-agent harness/model combinations specificallyLow
Humanity's Last ExamBroad expert-level knowledge (single eval, not composite)Low

Use the Intelligence Index for an at-a-glance overall capability ranking; use its constituent evaluations (or narrower benchmarks like HLE or SciCode) when you need to diagnose a specific capability gap.

Run the Artificial Analysis Intelligence Index on Your Model

Benchgen lets teams track composite intelligence-index-style performance across model versions, catching regressions in specific capability categories that a single top-line number would hide.

Frequently Asked Questions

What is the Artificial Analysis Intelligence Index? A composite score from Artificial Analysis combining 10 evaluations across agentic, coding, general-knowledge, and scientific-reasoning tasks into one weighted number for comparing overall model capability.
What does a good Intelligence Index score look like? As of the v4.1.1 methodology (September 2026), frontier models score in the 58-66 range, with no model above 66 — the composite is not yet close to saturated.
Who created the Artificial Analysis Intelligence Index? Artificial Analysis, an independent AI benchmarking organization. Full methodology is published at artificialanalysis.ai/methodology/intelligence-benchmarking.
Is the Artificial Analysis Intelligence Index saturated? No — scores remain well below the theoretical ceiling, and the composite's constituent evaluations are periodically refreshed (e.g. v4.1.1 to v4.3) specifically to keep headroom for differentiation as models improve.

Benchmark definition paraphrased from Artificial Analysis's published Intelligence Benchmarking methodology. State-of-the-art scores sourced from OpenAI's GPT-6 Astra announcement and attributed inline. Last updated 2026-09-07.