| Rank | Model | Score |
|---|---|---|
| 1 | qwen3-235b-a22b | 0.956 |
| 2 | qwen3-32b | 0.938 |
| 3 | qwen3-30b-a3b | 0.91 |
| 4 | llama-3-3-nemotron-super-49b-v1 | 0.883 |
| 5 | mistral-small-3-24b-instruct | 0.876 |
| 6 | qwen2-5-72b-instruct | 0.812 |
| 7 | phi-4-reasoning-plus | 0.79 |
| 8 | deepseek-v2-5 | 0.762 |
| 9 | phi-4 | 0.754 |
| 10 | phi-4-reasoning | 0.733 |
| 11 | ministral-8b-instruct | 0.709 |
| 12 | jamba-1-5-large | 0.654 |
| 13 | mistral-small-4 | 0.583 |
| 14 | granite-3-3-8b-base | 0.576 |
| 15 | granite-3-3-8b-instruct | 0.576 |
| 16 | ministral-3-14b-instruct-2512 | 0.551 |
| 17 | mistral-large-3 | 0.551 |
| 18 | qwen2-5-7b-instruct | 0.52 |
| 19 | ministral-3-8b-instruct-2512 | 0.509 |
| 20 | jamba-1-5-mini | 0.461 |
| 21 | mistral-small-3-2-24b-instruct | 0.431 |
| 22 | phi-3-5-moe-instruct | 0.379 |
| 23 | phi-3-5-mini-instruct | 0.37 |
| 24 | phi-4-mini | 0.328 |
| 25 | ministral-3-3b-instruct-2512 | 0.305 |
1 phaseActive
Automatic LLM benchmark using 500 challenging real-world prompts judged by frontier LLMs as a proxy for Chatbot Arena rankings.
Quick answer: Arena Hard (Li et al., 2024) is an automatic LLM evaluation framework using 500 challenging real-world user queries sourced from Chatbot Arena. Models are judged by frontier LLMs (GPT-4 family) acting as proxies for human preference, achieving 98.6% correlation with human preference rankings.
Arena Hard (Arena-Hard-Auto) is an automatic evaluation benchmark for instruction-tuned LLMs consisting of 500 challenging real-world prompts curated by BenchBuilder. It includes open-ended software engineering problems, mathematical questions, and creative writing tasks. The benchmark uses LLM-as-a-Judge methodology with frontier models as automatic judges to approximate human preference. Arena Hard achieves 98.6% correlation with human preference rankings and provides 3x higher model separation compared to MT-Bench.
| Property | Value |
|---|---|
| Tasks | 500 prompts |
| Metric | Win rate vs. baseline model |
| Categories | Reasoning, General, Creativity, Writing |
| Modality | Text |
| Judge | LLM-as-a-Judge (GPT-4 / Gemini) |
| Language | English |
Source: Li et al. 2024. Last updated 2026-07-24.