Benchgen
Models/mistral-ai/

Mistral Large 3

DraftPublic

Model Details

Mistral Large 3

Organization License Released

Quick answer: Mistral Large 3 is Mistral AI's June 2025 large model scoring 90.4% MATH, 55.1% Arena Hard, and 23.8% SimpleQA. Commercial tier — strong competition math.

At a Glance

Where Mistral Large 3 leads

  • 90.4% MATH — excellent competition math
  • Mistral Large tier: more capable than Small

Where it lags

  • 55.1% Arena Hard — surprisingly low (below Mistral Small 3 at 87.6%)
  • 23.8% SimpleQA — low factual accuracy
  • Commercial license: not Apache 2.0

Best for: Math-heavy Mistral API deployments where Large tier budget is available.

What Mistral Large 3 Is

Mistral Large 3 is the June 2025 "Large" tier API model from Mistral AI. The 90.4% MATH is strong, but the 55.1% Arena Hard is unexpectedly below Mistral Small 3 (87.6%), suggesting different optimisation targets.

Specifications

FieldValue
OrganizationMistral AI
LicenseCommercial (Mistral API)
Release dateJune 2025
ModalityText only

Pricing

Available via Mistral AI API. Refer to Mistral pricing.

Public Benchmark Scores

BenchmarkScoreSourceDate
MATH90.4%Benchgen evaluation2025-06
Arena Hard55.1%Benchgen evaluation2025-06
SimpleQA23.8%Benchgen evaluation2025-06

Mistral Large 3 vs Alternatives

ModelMATHArena HardLicense
Mistral Large 390.4%55.1%Commercial
Mistral Small 3 24B70.6%87.6%Apache 2.0
Mistral Large 2Mistral Research

Notably, Mistral Small 3 has higher Arena Hard (87.6% vs 55.1%). Use Large 3 specifically for MATH; use Small 3 for instruction-heavy tasks.

Frequently Asked Questions

What is Mistral Large 3? Mistral AI's June 2025 large tier scoring 90.4% MATH, 55.1% Arena Hard. Commercial API.

Specs from Mistral AI's Mistral Large 3 release (June 2025) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.