Quick answer: Mistral Small 3.2 24B Instruct is Mistral AI's June 2025 update scoring 43.1% Arena Hard, 69.4% MATH, 69.1% MMLU-Pro, and 12.1% SimpleQA. Apache 2.0.
Where Mistral Small 3.2 leads
Where it lags
Best for: Math-intensive pipelines needing Apache 2.0; use Mistral Small 3 for instruction quality; prefer 3.1 for Arena Hard tasks.
Mistral Small 3.2 24B Instruct is Mistral AI's June 2025 update to the Small 3 series. Notably, the Arena Hard drops significantly (43.1% vs 87.6% in Small 3) — suggesting this version trades instruction quality for other improvements.
For instruction-following tasks: use Mistral Small 3 (87.6% Arena Hard) or Mistral Small 3.1 (88.4% HumanEval). For math: 3.2 is comparable at 69.4%.
| Field | Value |
|---|---|
| Organization | Mistral AI |
| License | Apache 2.0 |
| HuggingFace | mistralai/Mistral-Small-3.2-24B-Instruct |
| Release date | June 2025 |
| Parameters | 24B |
| Modality | Text and vision |
Open weights under Apache 2.0. Also available via Mistral API.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| Arena Hard | 43.1% | Benchgen evaluation | 2025-06 |
| MATH | 69.4% | Benchgen evaluation | 2025-06 |
| MMLU-Pro | 69.1% | Benchgen evaluation | 2025-06 |
| SimpleQA | 12.1% | Benchgen evaluation | 2025-06 |
| Model | Arena Hard | MATH | MMLU-Pro | License |
|---|---|---|---|---|
| Mistral Small 3.2 24B | 43.1% | 69.4% | 69.1% | Apache 2.0 |
| Mistral Small 3 24B | 87.6% | 70.6% | 66.3% | Apache 2.0 |
| Mistral Small 3.1 24B | — | 69.3% | 66.8% | Apache 2.0 |
Mistral Small 3.2 has lower Arena Hard than 3.0 (43.1% vs 87.6%) — use Small 3 for instruction quality; use 3.2 only for multimodal or specific math tasks.
Specs from Mistral AI's Small 3.2 24B Instruct release (June 2025) and Benchgen evaluations. Last updated 2026-07-24.