Quick answer: Phi-4 Reasoning Plus is Microsoft's April 2025 open-weight reasoning variant of Phi-4, scoring 79% on Arena Hard, 53.1% on LiveCodeBench, and 76% on MMLU-Pro. It adds chain-of-thought reasoning to the base Phi-4 architecture with Apache 2.0 license — making it the strongest open reasoning model at the 14B scale.
Where Phi-4 Reasoning Plus leads
Where it lags
Best for: Open-weight reasoning at 14B scale; Apache 2.0 reasoning model for commercial deployments; STEM tasks requiring reasoning under GPU constraints.
Phi-4 Reasoning Plus is Microsoft's extension of Phi-4, adding explicit chain-of-thought reasoning training to the base 14B model. Released April 30, 2025, it was one of the first Apache 2.0 reasoning models at a deployable scale.
The "Plus" variant represents the most capable version of Phi-4 Reasoning — trained with additional reasoning data and reinforcement learning to strengthen its thinking capability. With 76% MMLU-Pro, it surpasses the base Phi-4 (70.4%) on academic knowledge when reasoning is enabled.
Phi-4 Reasoning Plus fills a distinctive niche: Apache 2.0 reasoning model under 20B parameters. For teams requiring a reasoning model that can be fine-tuned, self-hosted, and commercially deployed without proprietary restrictions, it is the primary open-weight option at this scale.
| Field | Value |
|---|---|
| Organization | Microsoft |
| Parameters | 14B (dense) |
| Context window | 32,000 tokens |
| License | Apache 2.0 |
| HuggingFace | microsoft/Phi-4-reasoning-plus |
| Release date | April 30, 2025 |
| Knowledge cutoff | September 2024 |
| Modality | Text only |
Phi-4 Reasoning Plus is available as open weights under Apache 2.0 — free for self-hosting and commercial use. Available via Azure AI Foundry and third-party providers.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| Arena Hard v2 | 79% | Benchgen evaluation | 2025-07 |
| LiveCodeBench | 53.1% | Benchgen evaluation | 2025-07 |
| MMLU-Pro | 76% | Benchgen evaluation | 2025-07 |
| Model | MMLU-Pro | LiveCodeBench | License | Reasoning |
|---|---|---|---|---|
| Phi-4 Reasoning Plus | 76% | 53.1% | Apache 2.0 | Yes |
| Phi-4 | 70.4% | — | Apache 2.0 | No |
| QwQ 32B | — | — | Apache 2.0 | Yes |
| Gemma 3 27B | 67.5% | — | Gemma ToU | No |
Phi-4 Reasoning Plus vs Phi-4: +5.6pp MMLU-Pro (76% vs 70.4%), stronger coding (53.1% LiveCodeBench). At 2× the latency due to reasoning steps. Choose Phi-4 for speed; Phi-4 Reasoning Plus for accuracy on complex tasks.
Specs from Microsoft's official Phi-4 Reasoning release (April 2025) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.