Quick answer: Phi-4 Reasoning is Microsoft's April 2025 reasoning-optimised model scoring 97.2% GSM8K, 92.7% MATH, 77.8% GPQA Diamond, 74.3% MMLU-Pro, 53.8% LiveCodeBench, and 73.3% Arena Hard. Apache 2.0.
Where Phi-4 Reasoning leads
Where it lags
Best for: Open-source math and reasoning tasks; Apache 2.0 reasoning-first pipelines; teams where Phi-4 Reasoning Plus is too large/expensive.
Phi-4 Reasoning is the chain-of-thought reasoning variant of Microsoft's Phi-4 (14B) model, released April 2025. It is the direct predecessor to Phi-4 Reasoning Plus, extending Phi-4's base capabilities with reasoning training (similar to o1-mini style chain-of-thought).
The 92.7% MATH and 77.8% GPQA Diamond scores show the reasoning training effect — substantially higher than Phi-4 Mini (64% MATH) and competitive with models several times larger.
| Field | Value |
|---|---|
| Organization | Microsoft |
| License | Apache 2.0 |
| HuggingFace | microsoft/Phi-4-reasoning |
| Release date | April 30, 2025 |
| Modality | Text only |
Open weights under Apache 2.0 — self-host at no cost. Available via Azure AI Foundry.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| GSM8K | 97.2% | Benchgen evaluation | 2025-04 |
| MATH | 92.7% | Benchgen evaluation | 2025-04 |
| GPQA Diamond | 77.8% | Benchgen evaluation | 2025-04 |
| MMLU-Pro | 74.3% | Benchgen evaluation | 2025-04 |
| LiveCodeBench | 53.8% | Benchgen evaluation | 2025-04 |
| Arena Hard | 73.3% | Benchgen evaluation | 2025-04 |
| Model | MATH | GPQA Diamond | MMLU-Pro | License |
|---|---|---|---|---|
| Phi-4 Reasoning | 92.7% | 77.8% | 74.3% | Apache 2.0 |
| Phi-4 Reasoning Plus | — | — | 76% | MIT |
| Phi-4 | — | — | — | MIT |
| QwQ-32B | — | — | — | Apache 2.0 |
Phi-4 Reasoning vs Phi-4 Reasoning Plus: Reasoning Plus scores 79% Arena Hard (vs 73.3%), 53.1% LiveCodeBench (vs 53.8%), 76% MMLU-Pro (vs 74.3%). Use Reasoning Plus for maximum Phi-4 reasoning capability; Reasoning for a slightly lighter variant.
Specs from Microsoft's Phi-4 Reasoning release (April 2025) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.