Quick answer: Mistral Small 3.1 24B Instruct is Mistral AI's March 2025 multimodal update scoring 88.4% HumanEval, 69.3% MATH, 66.8% MMLU-Pro, and 10.4% SimpleQA. Apache 2.0.
Where Mistral Small 3.1 leads
Where it lags
Best for: Apache 2.0 coding (HumanEval) with multimodal support; Small 3 series users needing vision capability.
Mistral Small 3.1 24B Instruct is the March 2025 update to Mistral Small 3, adding vision/multimodal support to the 24B model. The 88.4% HumanEval is a notable score — competitive with models twice its size.
| Field | Value |
|---|---|
| Organization | Mistral AI |
| License | Apache 2.0 |
| HuggingFace | mistralai/Mistral-Small-3.1-24B-Instruct |
| Release date | March 17, 2025 |
| Parameters | 24B |
| Modality | Text and vision (multimodal) |
| Context window | 128K tokens |
Open weights under Apache 2.0. Also available via Mistral API.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| HumanEval | 88.4% | Benchgen evaluation | 2025-03 |
| MATH | 69.3% | Benchgen evaluation | 2025-03 |
| MMLU-Pro | 66.8% | Benchgen evaluation | 2025-03 |
| SimpleQA | 10.4% | Benchgen evaluation | 2025-03 |
| Model | HumanEval | MATH | MMLU-Pro | Multimodal |
|---|---|---|---|---|
| Mistral Small 3.1 24B | 88.4% | 69.3% | 66.8% | Yes |
| Mistral Small 3 24B | 84.8% | 70.6% | 66.3% | No |
| Mistral Small 3.2 24B | — | 69.4% | 69.1% | Yes |
Mistral Small 3.1 vs 3.0: higher HumanEval (88.4% vs 84.8%), similar MATH. 3.1 adds multimodal. Choose 3.0 for Arena Hard instruction tasks.
Specs from Mistral AI's Small 3.1 24B Instruct release (March 2025) and Benchgen evaluations. Last updated 2026-07-24.