Benchgen

HealthBench — Results

RankModelScore
1kimi-k2-thinking-09050.58
2gpt-oss-120b-high0.576
3gpt-5-6-sol0.57
4gpt-5-6-terra0.57
5gpt-5-6-luna0.558
6gpt-5-3-codex0.541
7gpt-5-50.514

HealthBench

1 phaseActive

5,000-conversation healthcare benchmark graded by 262 physicians — measures AI performance and safety across health contexts and behavioral dimensions. Metric: rubric score.

Overview

HealthBench

Category Metric Saturation Tasks

Paper GitHub

Quick answer: HealthBench is an open-source healthcare benchmark by Arora et al. (2025) at OpenAI, consisting of 5,000 multi-turn conversations graded by 262 physicians using 48,562 unique rubric criteria. It measures both performance and safety of AI models in health contexts. Qwen3.8 Max leads with 60.2% across 9 evaluated models.


What Does HealthBench Test?

HealthBench evaluates AI models on realistic multi-turn medical conversations covering diverse health contexts and behavioral dimensions. The benchmark is physician-graded — each conversation is scored against rubric criteria written by 262 medical doctors, making it one of the most expert-validated healthcare benchmarks available.

DimensionWhat it tests
Medical accuracyCorrectness of health information provided
SafetyAvoidance of harmful advice
CommunicationClarity and appropriateness for patients
Clinician use casesProfessional-grade responses for practitioners

HealthBench also has two sub-benchmarks: HealthBench Consensus (highest physician agreement questions) and HealthBench Professional (clinician use cases).


How Is HealthBench Scored?

Each conversation is scored against physician-authored rubric criteria. Scores represent the fraction of rubric criteria met, normalized to 0–1. The overall score reflects both helpfulness and safety across the full conversation.


Key Facts

PropertyValue
PublishedMay 2025
Conversations5,000
Physician graders262
Rubric criteria48,562
MetricRubric score
Score range0–1
Top modelQwen3.8 Max (0.602)
Models evaluated9

FAQ

What is HealthBench? HealthBench is an open-source benchmark for measuring the performance and safety of large language models in healthcare, consisting of 5,000 multi-turn conversations evaluated by 262 physicians using 48,562 unique rubric criteria.

Who created HealthBench? HealthBench was created by Rahul K. Arora, Jason Wei, and colleagues at OpenAI, published in May 2025 (arXiv 2505.08775).

What are the HealthBench sub-benchmarks? HealthBench Consensus focuses on questions with especially high physician agreement. HealthBench Professional evaluates model capability for clinician use cases using real clinician-style chats.

What score does the best model achieve on HealthBench? Qwen3.8 Max from Alibaba Cloud currently leads with 0.602 (60.2%), followed by Kimi K2-Thinking-0905 at 0.580 and GPT OSS 120B at 0.576.