Quick answer: Mistral Large 2 is Mistral's July 2024 flagship model, scoring 93.0% on GSM8K, 92.0% on HumanEval, 84.0% on MMLU, and 8.63/10 on MT-Bench. With a 128K context window, it was Mistral's most capable model at launch and remains relevant for teams in the Mistral AI ecosystem.
Where Mistral Large 2 leads
Where it lags
Best for: Research and evaluation within the Mistral AI ecosystem; teams with Mistral Enterprise agreements; code generation tasks.
Mistral Large 2 (released July 24, 2024) was Mistral AI's highest-capability model in their "Large" tier, positioned above Mistral Large (2402) and below proprietary enterprise models. It offers strong coding (92.0% HumanEval) and reasoning (93.0% GSM8K) at 128K context.
The model is available on HuggingFace under the Mistral Research License — meaning the weights are accessible but commercial deployment requires a Mistral AI Enterprise agreement. For teams comparing open-weight options, Llama 3.3 70B (Llama 3.3 License) or DeepSeek-V3 (MIT) offer similar or better benchmarks with more permissive commercial terms.
| Field | Value |
|---|---|
| Organization | Mistral AI |
| Context window | 128,000 tokens |
| License | Mistral Research License |
| HuggingFace | mistralai/Mistral-Large-Instruct-2407 |
| Release date | July 24, 2024 |
| Knowledge cutoff | June 2024 |
| Modality | Text only |
Available via Mistral AI API (La Plateforme) with commercial pricing. Research use via HuggingFace under Mistral Research License.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| GSM8K | 93.0% | Benchgen evaluation | 2025-07 |
| HumanEval | 92.0% | Benchgen evaluation | 2025-07 |
| MMLU | 84.0% | Benchgen evaluation | 2025-07 |
| MT-Bench | 8.63 / 10 | Benchgen evaluation | 2025-07 |
| Model | GSM8K | HumanEval | MMLU | License |
|---|---|---|---|---|
| Mistral Large 2 | 93.0% | 92.0% | 84.0% | Research only |
| Llama 3.3 70B Instruct | — | 88.4% | 86.0% | Llama 3.3 |
| DeepSeek-V3 | — | — | 88.5% | MIT |
Mistral Large 2's coding scores (92% HumanEval) are competitive with Llama 3.3 70B (88.4%), but the Research License restricts commercial use. For permissive commercial deployment, Llama 3.3 70B or DeepSeek-V3 are preferred.
Specs from Mistral AI's official Mistral Large 2 release (July 2024) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.