Benchgen

AA-Omniscience — Results

RankModelScore
1claude-fable-540
2gemini-3-1-pro33
3gpt-5-6-sol22
4nemotron-3-5-lightning-30b-a3b16.6
5kimi-k2-66
6glm-5-24
7inkling2.1
8nemotron-3-ultra-550b-a55b-1
9kimi-k2-5-8
10deepseek-v4-pro-10
11lfm2-5-2-6b-29.5

AA-Omniscience

1 phaseActive

Artificial Analysis's knowledge and hallucination benchmark. 6,000 questions, 42 topics across 6 domains. Index -100 to +100 (rewards correct answers, penalizes hallucinations). Part of the AA Intelligence Index v4.1.

Overview

AA-Omniscience

Category Tasks Domains Range Saturation

Paper Dataset Leaderboard

Quick answer: AA-Omniscience is Artificial Analysis's benchmark for measuring factual recall and hallucination in LLMs. It contains 6,000 questions from authoritative sources across 6 major domains and 42 economically relevant topics. The AA-Omniscience Index (-100 to +100) rewards correct answers, penalizes hallucinations (wrong answers), and has no penalty for refusing to answer. A score of 0 means equal proportions of correct and incorrect answers. Most frontier models score below zero. Claude Fable 5 leads at 40. It's one of nine evaluations in the AA Intelligence Index v4.1.

At a Glance

What it tests: A model's ability to accurately recall factual knowledge across academic and professional domains while avoiding hallucination. The index simultaneously rewards accuracy and penalizes overconfidence — a model that always answers with false information will score -100; a model that always correctly answers or correctly abstains will score positively.

Why it matters: Standard accuracy metrics don't distinguish between "the model was wrong and guessed" versus "the model correctly abstained." AA-Omniscience's scoring structure incentivizes calibrated knowledge: models should answer when confident and abstain when uncertain. Most frontier models fail this calibration, scoring below zero — meaning they hallucinate more than they correctly answer on average.

Known limitations: Questions are generated by an LLM agent from authoritative sources (introducing possible generation artifacts). The benchmark is proprietary to Artificial Analysis, though a public subset is available on HuggingFace. The methodology details are published at artificialanalysis.ai/methodology/intelligence-benchmarking.

What AA-Omniscience Measures

AA-Omniscience contains 6,000 questions derived from authoritative academic and industry sources, spanning 42 economically relevant topics in 6 major domains:

DomainSelected Topics
BusinessAccounting, Corporate & Markets, Economics, Finance, Investments
Humanities & Social SciencesHistory, Law, Literature, Philosophy, Politics
Science, Engineering & MathBiology, Chemistry, Engineering, Mathematics, Physics
HealthBiomedical Sciences, Medicine, Public Health
LawConstitutional Law, Contract Law, Criminal Law, Tort Law
Software EngineeringPython, JavaScript/TypeScript, Go, Rust, SQL, and more

Scoring Formula

AA-Omniscience Index = (correct - incorrect) / total × 100

  • Correct answer: +1 point
  • Incorrect answer: -1 point
  • Abstaining / refusing: 0 points (no penalty)
  • Range: -100 to +100
  • Interpretation: 0 means equal numbers of correct and incorrect; positive = more correct than wrong; negative = more wrong than correct

Benchmark Specifications

FieldValue
Task categoryKnowledge / factuality / hallucination
MetricAA-Omniscience Index (-100 to +100)
Number of tasks6,000
Domains6 major domains, 42 topics
Questions derived fromAuthoritative academic and industry sources
Created byDeclan Jackson, William Keating, George Cameron, Micah Hill-Smith
AffiliationArtificial Analysis
PaperAA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models (arXiv 2511.13029)
DatasetArtificialAnalysis/AA-Omniscience-Public on HuggingFace
Leaderboardartificialanalysis.ai/evaluations/omniscience

State-of-the-Art Results

Scores from Inkling model card (Thinking Machines Lab, July 2026). All 9 comparison models reported.

RankModelIndexWeights
1Claude Fable 540Closed
2Gemini 3.1 Pro33Closed
3GPT-5.6 Sol22Closed
4Kimi K2.66Open
5GLM 5.24Open
6Inkling2.1Open
7Nemotron 3 Ultra 550B-1Open
8Kimi K2.5-8Open
9DeepSeek V4 Pro-10Open
BenchmarkTasksScoringPenalizes wrong?Saturation
AA-Omniscience6,000-100 to +100 indexYesLow
SimpleQA4,326% accuracyNoMedium
GPQA Diamond198% accuracyNoLow
MMLU-Pro12,000% accuracyNoMedium

Last updated 2026-07-16.