| Rank | Model | Score |
|---|---|---|
| 1 | mai-thinking-1 | 0.88 |
| 2 | phi-3-5-moe-instruct | 0.775 |
| 3 | granite-3-3-8b-instruct | 0.669 |
| 4 | phi-4-mini | 0.664 |
| 5 | phi-3-5-mini-instruct | 0.64 |
| 6 | hermes-3-70b | 0.633 |
| 7 | llama-3-1-nemotron-70b-instruct | 0.586 |
| 8 | qwen2-5-14b-instruct | 0.584 |
| 9 | jamba-1-5-large | 0.583 |
| 10 | granite-4-0-tiny-preview | 0.581 |
| 11 | qwen2-5-32b-instruct | 0.578 |
| 12 | command-r-plus | 0.563 |
| 13 | qwen2-72b-instruct | 0.548 |
| 14 | qwen2-5-coder-32b-instruct | 0.542 |
| 15 | jamba-1-5-mini | 0.541 |
| 16 | granite-3-3-8b-base | 0.521 |
| 17 | qwen2-5-coder-7b-instruct | 0.506 |
| 18 | mistral-nemo-instruct | 0.503 |
1 phaseActive
817-question benchmark measuring AI truthfulness on misconception-prone topics across health, law, finance, and politics. Metric: % truthful (MC).
Quick answer: TruthfulQA is a benchmark by Lin, Hilton, and Evans (2021) that measures whether language models generate truthful answers to 817 questions designed to elicit common human misconceptions. Questions span health, law, finance, politics, and 35 other categories. MAI-Thinking-1 leads with 88.0% across 18 evaluated models.
TruthfulQA tests AI models on questions where humans commonly give wrong answers due to false beliefs or misconceptions. A model trained on human text tends to pick up and repeat these falsehoods — TruthfulQA measures how well models resist this.
| Category | Examples |
|---|---|
| Health | Vaccine myths, medical misconceptions |
| Law | Legal misunderstandings, rights confusion |
| Finance | Investment myths, economic misconceptions |
| Politics | Electoral misconceptions, policy myths |
| History | Common historical falsehoods |
| Science | Popular science myths |
The benchmark is evaluated in multiple-choice (MC) format. Models are scored on the percentage of questions answered truthfully. A model that always picks the most statistically likely human answer would score poorly — the point is to measure resistance to common falsehoods.
| Property | Value |
|---|---|
| Published | September 2021 |
| Tasks | 817 questions |
| Categories | 38 |
| Metric | % truthful (MC) |
| Score range | 0–1 |
| Top model | MAI-Thinking-1 (0.880) |
| Models evaluated | 18 |
What is TruthfulQA? TruthfulQA measures whether AI language models generate truthful answers to questions where humans commonly make mistakes due to false beliefs or misconceptions, across topics including health, law, finance, and politics.
Who created TruthfulQA? TruthfulQA was created by Stephanie Lin, Jacob Hilton, and Owain Evans at the University of Oxford, published in September 2021 (arXiv 2109.07958).
Why do large models sometimes score low on TruthfulQA? Larger models trained on more human text can learn to reproduce common human falsehoods more fluently. The benchmark was specifically designed to surface this problem.
What score does the best model achieve on TruthfulQA? MAI-Thinking-1 from Microsoft achieves 0.880 (88.0%), followed by Phi-3.5-MoE-instruct at 0.775.