Quick answer: Phi-4 Mini is Microsoft's February 2025 compact model, scoring 88.6% GSM8K, 83.7% ARC-C, 64.0% MATH, and 52.8% MMLU-Pro. Apache 2.0 — Microsoft's compact model for edge and local deployment.
Where Phi-4 Mini leads
Where it lags
Best for: Edge deployment and local inference; mobile applications requiring math/reasoning; on-device code assistance.
Phi-4 Mini is Microsoft's February 2025 compact model — the small sibling of Phi-4 (14B). The Phi series is known for "small data, big training" philosophy: Microsoft invests heavily in curated, high-quality training data to achieve strong reasoning in small models.
Phi-4 Mini follows this pattern: 88.6% GSM8K at compact size is strong for the era, reflecting focused math training. The 64.0% MATH score is particularly high for a mini-class model. Apache 2.0 makes it freely deployable for commercial use.
| Field | Value |
|---|---|
| Organization | Microsoft |
| License | Apache 2.0 |
| HuggingFace | microsoft/Phi-4-mini-instruct |
| Release date | February 4, 2025 |
| Modality | Text only |
Open weights under Apache 2.0 — self-host at no cost. Available via Azure AI Foundry and Ollama/local inference.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| GSM8K | 88.6% | Benchgen evaluation | 2025-02 |
| ARC-C | 83.7% | Benchgen evaluation | 2025-02 |
| MATH | 64.0% | Benchgen evaluation | 2025-02 |
| MMLU-Pro | 52.8% | Benchgen evaluation | 2025-02 |
| HellaSwag | 69.1% | Benchgen evaluation | 2025-02 |
| Arena Hard | 32.8% | Benchgen evaluation | 2025-02 |
| Model | GSM8K | MATH | License | Size |
|---|---|---|---|---|
| Phi-4 Mini | 88.6% | 64.0% | Apache 2.0 | Mini |
| Phi-4 | — | — | MIT | 14B |
| Phi-4 Reasoning Plus | — | — | MIT | 14B |
| Llama 3.2 3B Instruct | 77.7% | — | Llama 3.2 | 3B |
Phi-4 Mini vs Llama 3.2 3B: higher GSM8K (88.6% vs 77.7%) and MATH (64%). For maximum compact math/reasoning under Apache 2.0: Phi-4 Mini. For Llama ecosystem compatibility: Llama 3.2 3B.
Specs from Microsoft's Phi-4 Mini release (February 2025) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.