| Rank | Model | Score |
|---|---|---|
| 1 | claude-fable-5 | 40 |
| 2 | gemini-3-1-pro | 33 |
| 3 | gpt-5-6-sol | 22 |
| 4 | nemotron-3-5-lightning-30b-a3b | 16.6 |
| 5 | kimi-k2-6 | 6 |
| 6 | glm-5-2 | 4 |
| 7 | inkling | 2.1 |
| 8 | nemotron-3-ultra-550b-a55b | -1 |
| 9 | kimi-k2-5 | -8 |
| 10 | deepseek-v4-pro | -10 |
| 11 | lfm2-5-2-6b | -29.5 |
1 phaseActive
Artificial Analysis's knowledge and hallucination benchmark. 6,000 questions, 42 topics across 6 domains. Index -100 to +100 (rewards correct answers, penalizes hallucinations). Part of the AA Intelligence Index v4.1.
Quick answer: AA-Omniscience is Artificial Analysis's benchmark for measuring factual recall and hallucination in LLMs. It contains 6,000 questions from authoritative sources across 6 major domains and 42 economically relevant topics. The AA-Omniscience Index (-100 to +100) rewards correct answers, penalizes hallucinations (wrong answers), and has no penalty for refusing to answer. A score of 0 means equal proportions of correct and incorrect answers. Most frontier models score below zero. Claude Fable 5 leads at 40. It's one of nine evaluations in the AA Intelligence Index v4.1.
What it tests: A model's ability to accurately recall factual knowledge across academic and professional domains while avoiding hallucination. The index simultaneously rewards accuracy and penalizes overconfidence — a model that always answers with false information will score -100; a model that always correctly answers or correctly abstains will score positively.
Why it matters: Standard accuracy metrics don't distinguish between "the model was wrong and guessed" versus "the model correctly abstained." AA-Omniscience's scoring structure incentivizes calibrated knowledge: models should answer when confident and abstain when uncertain. Most frontier models fail this calibration, scoring below zero — meaning they hallucinate more than they correctly answer on average.
Known limitations: Questions are generated by an LLM agent from authoritative sources (introducing possible generation artifacts). The benchmark is proprietary to Artificial Analysis, though a public subset is available on HuggingFace. The methodology details are published at artificialanalysis.ai/methodology/intelligence-benchmarking.
AA-Omniscience contains 6,000 questions derived from authoritative academic and industry sources, spanning 42 economically relevant topics in 6 major domains:
| Domain | Selected Topics |
|---|---|
| Business | Accounting, Corporate & Markets, Economics, Finance, Investments |
| Humanities & Social Sciences | History, Law, Literature, Philosophy, Politics |
| Science, Engineering & Math | Biology, Chemistry, Engineering, Mathematics, Physics |
| Health | Biomedical Sciences, Medicine, Public Health |
| Law | Constitutional Law, Contract Law, Criminal Law, Tort Law |
| Software Engineering | Python, JavaScript/TypeScript, Go, Rust, SQL, and more |
AA-Omniscience Index = (correct - incorrect) / total × 100
| Field | Value |
|---|---|
| Task category | Knowledge / factuality / hallucination |
| Metric | AA-Omniscience Index (-100 to +100) |
| Number of tasks | 6,000 |
| Domains | 6 major domains, 42 topics |
| Questions derived from | Authoritative academic and industry sources |
| Created by | Declan Jackson, William Keating, George Cameron, Micah Hill-Smith |
| Affiliation | Artificial Analysis |
| Paper | AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models (arXiv 2511.13029) |
| Dataset | ArtificialAnalysis/AA-Omniscience-Public on HuggingFace |
| Leaderboard | artificialanalysis.ai/evaluations/omniscience |
Scores from Inkling model card (Thinking Machines Lab, July 2026). All 9 comparison models reported.
| Rank | Model | Index | Weights |
|---|---|---|---|
| 1 | Claude Fable 5 | 40 | Closed |
| 2 | Gemini 3.1 Pro | 33 | Closed |
| 3 | GPT-5.6 Sol | 22 | Closed |
| 4 | Kimi K2.6 | 6 | Open |
| 5 | GLM 5.2 | 4 | Open |
| 6 | Inkling | 2.1 | Open |
| 7 | Nemotron 3 Ultra 550B | -1 | Open |
| 8 | Kimi K2.5 | -8 | Open |
| 9 | DeepSeek V4 Pro | -10 | Open |
| Benchmark | Tasks | Scoring | Penalizes wrong? | Saturation |
|---|---|---|---|---|
| AA-Omniscience | 6,000 | -100 to +100 index | Yes | Low |
| SimpleQA | 4,326 | % accuracy | No | Medium |
| GPQA Diamond | 198 | % accuracy | No | Low |
| MMLU-Pro | 12,000 | % accuracy | No | Medium |
Last updated 2026-07-16.