Benchgen
Models/microsoft/

Phi 4 Mini

DraftPublic

Model Details

Phi-4 Mini

Organization License Released

Quick answer: Phi-4 Mini is Microsoft's February 2025 compact model, scoring 88.6% GSM8K, 83.7% ARC-C, 64.0% MATH, and 52.8% MMLU-Pro. Apache 2.0 — Microsoft's compact model for edge and local deployment.

At a Glance

Where Phi-4 Mini leads

  • 88.6% GSM8K — excellent math for a compact model
  • 83.7% ARC-C — strong commonsense reasoning
  • 64.0% MATH — solid competition math
  • Apache 2.0 — fully open
  • Optimised for low-resource deployment (mobile, edge)

Where it lags

  • 32.8% Arena Hard — lower instruction following quality
  • 52.8% MMLU-Pro — limited academic breadth
  • 69.1% HellaSwag — moderate language understanding
  • Small model: limited capacity vs Phi-4 (14B)

Best for: Edge deployment and local inference; mobile applications requiring math/reasoning; on-device code assistance.

What Phi-4 Mini Is

Phi-4 Mini is Microsoft's February 2025 compact model — the small sibling of Phi-4 (14B). The Phi series is known for "small data, big training" philosophy: Microsoft invests heavily in curated, high-quality training data to achieve strong reasoning in small models.

Phi-4 Mini follows this pattern: 88.6% GSM8K at compact size is strong for the era, reflecting focused math training. The 64.0% MATH score is particularly high for a mini-class model. Apache 2.0 makes it freely deployable for commercial use.

Specifications

FieldValue
OrganizationMicrosoft
LicenseApache 2.0
HuggingFacemicrosoft/Phi-4-mini-instruct
Release dateFebruary 4, 2025
ModalityText only

Pricing

Open weights under Apache 2.0 — self-host at no cost. Available via Azure AI Foundry and Ollama/local inference.

Public Benchmark Scores

BenchmarkScoreSourceDate
GSM8K88.6%Benchgen evaluation2025-02
ARC-C83.7%Benchgen evaluation2025-02
MATH64.0%Benchgen evaluation2025-02
MMLU-Pro52.8%Benchgen evaluation2025-02
HellaSwag69.1%Benchgen evaluation2025-02
Arena Hard32.8%Benchgen evaluation2025-02

Phi-4 Mini vs Alternatives

ModelGSM8KMATHLicenseSize
Phi-4 Mini88.6%64.0%Apache 2.0Mini
Phi-4MIT14B
Phi-4 Reasoning PlusMIT14B
Llama 3.2 3B Instruct77.7%Llama 3.23B

Phi-4 Mini vs Llama 3.2 3B: higher GSM8K (88.6% vs 77.7%) and MATH (64%). For maximum compact math/reasoning under Apache 2.0: Phi-4 Mini. For Llama ecosystem compatibility: Llama 3.2 3B.

Frequently Asked Questions

What is Phi-4 Mini? Microsoft's February 2025 compact model scoring 88.6% GSM8K, 83.7% ARC-C, 64.0% MATH, and 52.8% MMLU-Pro. Apache 2.0 for edge and local deployment.
Should I use Phi-4 Mini or Phi-4? Phi-4 (14B) for maximum capability within the Phi family. Phi-4 Mini for edge devices, mobile, or deployments where model size is the primary constraint.

Specs from Microsoft's Phi-4 Mini release (February 2025) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.