Benchgen

Arena Hard — Results

RankModelScore
1qwen3-235b-a22b0.956
2qwen3-32b0.938
3qwen3-30b-a3b0.91
4llama-3-3-nemotron-super-49b-v10.883
5mistral-small-3-24b-instruct0.876
6qwen2-5-72b-instruct0.812
7phi-4-reasoning-plus0.79
8deepseek-v2-50.762
9phi-40.754
10phi-4-reasoning0.733
11ministral-8b-instruct0.709
12jamba-1-5-large0.654
13mistral-small-40.583
14granite-3-3-8b-base0.576
15granite-3-3-8b-instruct0.576
16ministral-3-14b-instruct-25120.551
17mistral-large-30.551
18qwen2-5-7b-instruct0.52
19ministral-3-8b-instruct-25120.509
20jamba-1-5-mini0.461
21mistral-small-3-2-24b-instruct0.431
22phi-3-5-moe-instruct0.379
23phi-3-5-mini-instruct0.37
24phi-4-mini0.328
25ministral-3-3b-instruct-25120.305
A

Arena Hard

1 phaseActive

Automatic LLM benchmark using 500 challenging real-world prompts judged by frontier LLMs as a proxy for Chatbot Arena rankings.

Overview

Arena Hard

Category Metric Tasks Saturation Created

Paper GitHub Dataset

Quick answer: Arena Hard (Li et al., 2024) is an automatic LLM evaluation framework using 500 challenging real-world user queries sourced from Chatbot Arena. Models are judged by frontier LLMs (GPT-4 family) acting as proxies for human preference, achieving 98.6% correlation with human preference rankings.

What is Arena Hard?

Arena Hard (Arena-Hard-Auto) is an automatic evaluation benchmark for instruction-tuned LLMs consisting of 500 challenging real-world prompts curated by BenchBuilder. It includes open-ended software engineering problems, mathematical questions, and creative writing tasks. The benchmark uses LLM-as-a-Judge methodology with frontier models as automatic judges to approximate human preference. Arena Hard achieves 98.6% correlation with human preference rankings and provides 3x higher model separation compared to MT-Bench.

Benchmark Details

PropertyValue
Tasks500 prompts
MetricWin rate vs. baseline model
CategoriesReasoning, General, Creativity, Writing
ModalityText
JudgeLLM-as-a-Judge (GPT-4 / Gemini)
LanguageEnglish

Source: Li et al. 2024. Last updated 2026-07-24.