Benchgen

TruthfulQA — Results

RankModelScore
1mai-thinking-10.88
2phi-3-5-moe-instruct0.775
3granite-3-3-8b-instruct0.669
4phi-4-mini0.664
5phi-3-5-mini-instruct0.64
6hermes-3-70b0.633
7llama-3-1-nemotron-70b-instruct0.586
8qwen2-5-14b-instruct0.584
9jamba-1-5-large0.583
10granite-4-0-tiny-preview0.581
11qwen2-5-32b-instruct0.578
12command-r-plus0.563
13qwen2-72b-instruct0.548
14qwen2-5-coder-32b-instruct0.542
15jamba-1-5-mini0.541
16granite-3-3-8b-base0.521
17qwen2-5-coder-7b-instruct0.506
18mistral-nemo-instruct0.503

TruthfulQA

1 phaseActive

817-question benchmark measuring AI truthfulness on misconception-prone topics across health, law, finance, and politics. Metric: % truthful (MC).

Overview

TruthfulQA

Category Metric Saturation Tasks

Paper GitHub Dataset

Quick answer: TruthfulQA is a benchmark by Lin, Hilton, and Evans (2021) that measures whether language models generate truthful answers to 817 questions designed to elicit common human misconceptions. Questions span health, law, finance, politics, and 35 other categories. MAI-Thinking-1 leads with 88.0% across 18 evaluated models.


What Does TruthfulQA Test?

TruthfulQA tests AI models on questions where humans commonly give wrong answers due to false beliefs or misconceptions. A model trained on human text tends to pick up and repeat these falsehoods — TruthfulQA measures how well models resist this.

CategoryExamples
HealthVaccine myths, medical misconceptions
LawLegal misunderstandings, rights confusion
FinanceInvestment myths, economic misconceptions
PoliticsElectoral misconceptions, policy myths
HistoryCommon historical falsehoods
SciencePopular science myths

How Is TruthfulQA Scored?

The benchmark is evaluated in multiple-choice (MC) format. Models are scored on the percentage of questions answered truthfully. A model that always picks the most statistically likely human answer would score poorly — the point is to measure resistance to common falsehoods.


Key Facts

PropertyValue
PublishedSeptember 2021
Tasks817 questions
Categories38
Metric% truthful (MC)
Score range0–1
Top modelMAI-Thinking-1 (0.880)
Models evaluated18

FAQ

What is TruthfulQA? TruthfulQA measures whether AI language models generate truthful answers to questions where humans commonly make mistakes due to false beliefs or misconceptions, across topics including health, law, finance, and politics.

Who created TruthfulQA? TruthfulQA was created by Stephanie Lin, Jacob Hilton, and Owain Evans at the University of Oxford, published in September 2021 (arXiv 2109.07958).

Why do large models sometimes score low on TruthfulQA? Larger models trained on more human text can learn to reproduce common human falsehoods more fluently. The benchmark was specifically designed to surface this problem.

What score does the best model achieve on TruthfulQA? MAI-Thinking-1 from Microsoft achieves 0.880 (88.0%), followed by Phi-3.5-MoE-instruct at 0.775.