Quick answer: Mistral Small 3 24B Instruct is Mistral AI's January 2025 Apache 2.0 model scoring 87.6% Arena Hard, 84.8% HumanEval, 70.6% MATH, 66.3% MMLU-Pro, and 8.35 MT-Bench.
Where Mistral Small 3 leads
Where it lags
Best for: Instruction-following pipelines needing high Arena Hard quality; Apache 2.0 deployment at 24B; multi-turn chat and coding.
Mistral Small 3 24B Instruct is Mistral AI's January 2025 "Small" model release — a 24B Apache 2.0 model optimised for instruction following and coding. The 87.6% Arena Hard is the standout metric — one of the highest for any 24B open model at this period.
| Field | Value |
|---|---|
| Organization | Mistral AI |
| License | Apache 2.0 |
| HuggingFace | mistralai/Mistral-Small-3.1-24B-Instruct |
| Release date | January 30, 2025 |
| Parameters | 24B |
| Modality | Text only |
| Context window | 128K tokens |
Open weights under Apache 2.0. Also available via Mistral API.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| Arena Hard | 87.6% | Benchgen evaluation | 2025-01 |
| HumanEval | 84.8% | Benchgen evaluation | 2025-01 |
| MATH | 70.6% | Benchgen evaluation | 2025-01 |
| MMLU-Pro | 66.3% | Benchgen evaluation | 2025-01 |
| MT-Bench | 8.35 | Benchgen evaluation | 2025-01 |
| Model | Arena Hard | HumanEval | MMLU-Pro | License |
|---|---|---|---|---|
| Mistral Small 3 24B | 87.6% | 84.8% | 66.3% | Apache 2.0 |
| Qwen2.5 14B Instruct | — | 83.5% | 64.0% | Apache 2.0 |
| Mistral Large 2 | — | — | — | Mistral Research |
Mistral Small 3 vs Qwen2.5 14B: higher HumanEval (84.8% vs 83.5%), slightly higher MMLU-Pro (66.3% vs 64.0%), plus Arena Hard coverage. For high-Arena-Hard at Apache 2.0: Mistral Small 3.
Specs from Mistral AI's Mistral Small 3 24B Instruct release (January 2025) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.