Benchgen
Models/microsoft/

Phi 4 Reasoning

DraftPublic

Model Details

Phi-4 Reasoning

Organization License Released Reasoning

Quick answer: Phi-4 Reasoning is Microsoft's April 2025 reasoning-optimised model scoring 97.2% GSM8K, 92.7% MATH, 77.8% GPQA Diamond, 74.3% MMLU-Pro, 53.8% LiveCodeBench, and 73.3% Arena Hard. Apache 2.0.

At a Glance

Where Phi-4 Reasoning leads

  • 97.2% GSM8K — near-perfect grade-school math
  • 92.7% MATH — excellent competition math
  • 77.8% GPQA Diamond — solid graduate-level science
  • 53.8% LiveCodeBench — competitive coding
  • Apache 2.0 — fully open

Where it lags

  • 73.3% Arena Hard — moderate instruction quality
  • 74.3% MMLU-Pro — moderate academic breadth
  • Superseded by Phi-4 Reasoning Plus on most metrics

Best for: Open-source math and reasoning tasks; Apache 2.0 reasoning-first pipelines; teams where Phi-4 Reasoning Plus is too large/expensive.

What Phi-4 Reasoning Is

Phi-4 Reasoning is the chain-of-thought reasoning variant of Microsoft's Phi-4 (14B) model, released April 2025. It is the direct predecessor to Phi-4 Reasoning Plus, extending Phi-4's base capabilities with reasoning training (similar to o1-mini style chain-of-thought).

The 92.7% MATH and 77.8% GPQA Diamond scores show the reasoning training effect — substantially higher than Phi-4 Mini (64% MATH) and competitive with models several times larger.

Specifications

FieldValue
OrganizationMicrosoft
LicenseApache 2.0
HuggingFacemicrosoft/Phi-4-reasoning
Release dateApril 30, 2025
ModalityText only

Pricing

Open weights under Apache 2.0 — self-host at no cost. Available via Azure AI Foundry.

Public Benchmark Scores

BenchmarkScoreSourceDate
GSM8K97.2%Benchgen evaluation2025-04
MATH92.7%Benchgen evaluation2025-04
GPQA Diamond77.8%Benchgen evaluation2025-04
MMLU-Pro74.3%Benchgen evaluation2025-04
LiveCodeBench53.8%Benchgen evaluation2025-04
Arena Hard73.3%Benchgen evaluation2025-04

Phi-4 Reasoning vs Alternatives

ModelMATHGPQA DiamondMMLU-ProLicense
Phi-4 Reasoning92.7%77.8%74.3%Apache 2.0
Phi-4 Reasoning Plus76%MIT
Phi-4MIT
QwQ-32BApache 2.0

Phi-4 Reasoning vs Phi-4 Reasoning Plus: Reasoning Plus scores 79% Arena Hard (vs 73.3%), 53.1% LiveCodeBench (vs 53.8%), 76% MMLU-Pro (vs 74.3%). Use Reasoning Plus for maximum Phi-4 reasoning capability; Reasoning for a slightly lighter variant.

Frequently Asked Questions

What is Phi-4 Reasoning? Microsoft's April 2025 reasoning-optimised 14B model scoring 97.2% GSM8K, 92.7% MATH, 77.8% GPQA Diamond. Apache 2.0.
What is the difference between Phi-4 Reasoning and Phi-4 Reasoning Plus? Phi-4 Reasoning Plus is a stronger variant with higher Arena Hard (79% vs 73.3%) and MMLU-Pro (76% vs 74.3%). Phi-4 Reasoning is the base reasoning variant. Both are Microsoft-built on the Phi-4 architecture. Reasoning Plus is the recommended choice.

Specs from Microsoft's Phi-4 Reasoning release (April 2025) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.