| Rank | Model | Score |
|---|---|---|
| 1 | gpt-6-astra | 61.2 |
1 phaseActive
Artificial Analysis's composite index across agents, coding, general knowledge, and scientific reasoning — weighted average of 10 evaluations.
Quick answer: The Artificial Analysis Intelligence Index is a composite score combining 10 underlying evaluations — spanning agentic tasks (30% weight), coding (20%), scientific reasoning (20%), and general knowledge (30%) — into a single number for comparing overall model capability. As of v4.3, the suite includes AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, AA-Omniscience, GDP.pdf, AA-LCR v1.1, Humanity's Last Exam, and CritPt.
What it tests: Overall model capability across four weighted categories — Agents (30%), Coding (20%), Scientific Reasoning (20%), General (30%) — synthesized from 10 individual evaluation datasets into one comparable number.
Why it matters: No single benchmark captures overall model quality; the Index is designed as the best available one-number synthesis for cross-model comparison, deliberately weighted toward agentic capability (30%) to reflect real-world deployment patterns.
Known limitations: As a composite, the Index can mask category-specific weaknesses — a model can score well overall while lagging badly in one category. It's also version-specific (methodology and constituent evaluations change between versions, e.g. v4.1.1 vs v4.3), so scores are only directly comparable within the same Index version.
The Index combines a suite of evaluation datasets designed to assess language model capabilities across reasoning, knowledge, math, and programming into one weighted composite score. Artificial Analysis positions it as a more useful cross-model comparison than any single existing metric, while being explicit that — like all such syntheses — it has limitations and may not map directly onto every use case.
As of v4.3, the Index incorporates 10 evaluations across four categories: Agents (30% total weight: AA-Briefcase 15%, GDPval-AA v2 10%, AutomationBench-AA 5%), Coding (20%: Terminal-Bench v4.0 10%, SciCode 10%), General (30%: AA-Omniscience 15%, GDP.pdf 10%, AA-LCR v1.1 5%), and Scientific Reasoning (20%: Humanity's Last Exam 10%, CritPt 10%). Artificial Analysis reports a 95% confidence interval of under ±1% for the composite, based on repeated-run experiments, though individual constituent evaluations may have wider confidence intervals.
It's a primarily text-based, English-language evaluation suite — image inputs, speech inputs, and multilingual performance are benchmarked separately and are not part of the Index score itself.
| Field | Value |
|---|---|
| Task category | Composite / Reasoning |
| Metric | Weighted-average composite index score |
| Category weights | Agents 30%, General 30%, Coding 20%, Scientific Reasoning 20% |
| Constituent evaluations (v4.3) | AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, AA-Omniscience, GDP.pdf, AA-LCR v1.1, Humanity's Last Exam, CritPt |
| Saturation | Low — frontier models still in the 50-65 range |
| Created by | Artificial Analysis |
| Methodology | artificialanalysis.ai/methodology/intelligence-benchmarking |
Each constituent evaluation produces its own score (accuracy, pass@1, Elo-based pairwise comparison, etc., depending on the eval), which Artificial Analysis then combines into a single weighted-average composite using the category weights above. Scores are typically reported on a scale similar to a percentage (roughly 0-100), though the composite is best read as a relative ranking tool rather than an absolute accuracy figure, since it blends fundamentally different scoring types (accuracy, Elo, pass-rate) into one number.
| Rank | Model | Score | Version | Source | Date |
|---|---|---|---|---|---|
| 1 | Claude Fable 5.1 | 65.7 | v4.1.1 | OpenAI: GPT-6 Astra | 2026-09 |
| 2 | Claude Opus 5 | 63.1 | v4.1.1 | OpenAI: GPT-6 Astra | 2026-09 |
| 3 | GPT-6 Astra | 61.2 | v4.1.1 | OpenAI: GPT-6 Astra | 2026-09 |
| 4 | Claude Fable 5 | 62.1 | v4.1.1 | OpenAI: GPT-6 Astra | 2026-09 |
| 5 | GPT-5.6 Sol | 60.9 | v4.1.1 | OpenAI: GPT-6 Astra | 2026-09 |
| 6 | Gemini 3 Flash | 58.7 | v4.1.1 | OpenAI: GPT-6 Astra | 2026-09 |
Scores above are on Index v4.1.1 as reported in OpenAI's GPT-6 Astra announcement (September 2026) — note this predates the v4.3 methodology described in Artificial Analysis's own published methodology (10-evaluation set with AA-Briefcase/GDP.pdf added). Scores across different Index versions are not directly comparable.
No Benchgen results yet — be the first to run the Artificial Analysis Intelligence Index suite.
| Benchmark | What it tests | Saturation |
|---|---|---|
| AA Intelligence Index | Composite: agents, coding, general, scientific reasoning | Low |
| AA Coding Agent Index | Coding-agent harness/model combinations specifically | Low |
| Humanity's Last Exam | Broad expert-level knowledge (single eval, not composite) | Low |
Use the Intelligence Index for an at-a-glance overall capability ranking; use its constituent evaluations (or narrower benchmarks like HLE or SciCode) when you need to diagnose a specific capability gap.
Benchgen lets teams track composite intelligence-index-style performance across model versions, catching regressions in specific capability categories that a single top-line number would hide.
Benchmark definition paraphrased from Artificial Analysis's published Intelligence Benchmarking methodology. State-of-the-art scores sourced from OpenAI's GPT-6 Astra announcement and attributed inline. Last updated 2026-09-07.