Benchgen
Models/mistral-ai/

Mistral Small 3 24B Instruct

DraftPublic

Model Details

Mistral Small 3 24B Instruct

Organization License Released Params

Quick answer: Mistral Small 3 24B Instruct is Mistral AI's January 2025 Apache 2.0 model scoring 87.6% Arena Hard, 84.8% HumanEval, 70.6% MATH, 66.3% MMLU-Pro, and 8.35 MT-Bench.

At a Glance

Where Mistral Small 3 leads

  • 87.6% Arena Hard — strong instruction quality, leading among 24B models
  • 84.8% HumanEval — competitive coding
  • 8.35 MT-Bench — excellent multi-turn conversation
  • Apache 2.0 — fully open
  • 24B: smaller than Mistral Large 2 but significantly stronger quality

Where it lags

  • 66.3% MMLU-Pro — moderate advanced knowledge
  • 70.6% MATH — moderate competition math

Best for: Instruction-following pipelines needing high Arena Hard quality; Apache 2.0 deployment at 24B; multi-turn chat and coding.

What Mistral Small 3 24B Is

Mistral Small 3 24B Instruct is Mistral AI's January 2025 "Small" model release — a 24B Apache 2.0 model optimised for instruction following and coding. The 87.6% Arena Hard is the standout metric — one of the highest for any 24B open model at this period.

Specifications

FieldValue
OrganizationMistral AI
LicenseApache 2.0
HuggingFacemistralai/Mistral-Small-3.1-24B-Instruct
Release dateJanuary 30, 2025
Parameters24B
ModalityText only
Context window128K tokens

Pricing

Open weights under Apache 2.0. Also available via Mistral API.

Public Benchmark Scores

BenchmarkScoreSourceDate
Arena Hard87.6%Benchgen evaluation2025-01
HumanEval84.8%Benchgen evaluation2025-01
MATH70.6%Benchgen evaluation2025-01
MMLU-Pro66.3%Benchgen evaluation2025-01
MT-Bench8.35Benchgen evaluation2025-01

Mistral Small 3 vs Alternatives

ModelArena HardHumanEvalMMLU-ProLicense
Mistral Small 3 24B87.6%84.8%66.3%Apache 2.0
Qwen2.5 14B Instruct83.5%64.0%Apache 2.0
Mistral Large 2Mistral Research

Mistral Small 3 vs Qwen2.5 14B: higher HumanEval (84.8% vs 83.5%), slightly higher MMLU-Pro (66.3% vs 64.0%), plus Arena Hard coverage. For high-Arena-Hard at Apache 2.0: Mistral Small 3.

Frequently Asked Questions

What is Mistral Small 3 24B Instruct? Mistral AI's January 2025 24B model scoring 87.6% Arena Hard, 84.8% HumanEval, 66.3% MMLU-Pro. Apache 2.0 — top Arena Hard at 24B scale.

Specs from Mistral AI's Mistral Small 3 24B Instruct release (January 2025) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.