Quick answer: Mistral Large 3 is Mistral AI's June 2025 large model scoring 90.4% MATH, 55.1% Arena Hard, and 23.8% SimpleQA. Commercial tier — strong competition math.
Where Mistral Large 3 leads
Where it lags
Best for: Math-heavy Mistral API deployments where Large tier budget is available.
Mistral Large 3 is the June 2025 "Large" tier API model from Mistral AI. The 90.4% MATH is strong, but the 55.1% Arena Hard is unexpectedly below Mistral Small 3 (87.6%), suggesting different optimisation targets.
| Field | Value |
|---|---|
| Organization | Mistral AI |
| License | Commercial (Mistral API) |
| Release date | June 2025 |
| Modality | Text only |
Available via Mistral AI API. Refer to Mistral pricing.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| MATH | 90.4% | Benchgen evaluation | 2025-06 |
| Arena Hard | 55.1% | Benchgen evaluation | 2025-06 |
| SimpleQA | 23.8% | Benchgen evaluation | 2025-06 |
| Model | MATH | Arena Hard | License |
|---|---|---|---|
| Mistral Large 3 | 90.4% | 55.1% | Commercial |
| Mistral Small 3 24B | 70.6% | 87.6% | Apache 2.0 |
| Mistral Large 2 | — | — | Mistral Research |
Notably, Mistral Small 3 has higher Arena Hard (87.6% vs 55.1%). Use Large 3 specifically for MATH; use Small 3 for instruction-heavy tasks.
Specs from Mistral AI's Mistral Large 3 release (June 2025) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.