Benchgen
Models/microsoft/

Phi-3.5-MoE-instruct

DraftPublic

Model Details

Phi-3.5 MoE Instruct

Organization License Released MoE

Quick answer: Phi-3.5 MoE Instruct is Microsoft's August 2024 MoE instruction model scoring 91.0% ARC-C, 88.7% GSM8K, 78.5% MMLU, 59.5% MATH, and 45.3% MMLU-Pro. Apache 2.0.

At a Glance

Where Phi-3.5 MoE Instruct leads

  • 91.0% ARC-C — excellent common-sense reasoning
  • 88.7% GSM8K — solid grade-school math
  • 78.5% MMLU — strong general knowledge
  • Apache 2.0 — fully open
  • MoE efficiency: high capacity at low active-param cost

Where it lags

  • 59.5% MATH — moderate competition math
  • 45.3% MMLU-Pro — below average for pro-level knowledge
  • 37.9% Arena Hard — weaker instruction following
  • August 2024: older than Phi-4 / Phi-4 Reasoning

Best for: Open Apache 2.0 MoE deployments on MMLU/ARC/GSM8K tasks; legacy Phi-3.5 ecosystem; cost-efficient knowledge retrieval.

What Phi-3.5 MoE Instruct Is

Phi-3.5 MoE Instruct is Microsoft's August 2024 mixture-of-experts instruction model in the Phi-3.5 family. It combines Phi-3's efficient training with a MoE architecture, offering competitive knowledge benchmark scores at reduced per-token compute.

It was superseded by Phi-4 (November 2024) and Phi-4 Reasoning (April 2025), both of which substantially improve on MATH and MMLU-Pro.

Specifications

FieldValue
OrganizationMicrosoft
LicenseApache 2.0
HuggingFacemicrosoft/Phi-3.5-MoE-instruct
Release dateAugust 20, 2024
ArchitectureMoE
ModalityText only
Context window128K tokens

Pricing

Open weights under Apache 2.0 — self-host at no cost. Available via Azure AI Foundry.

Public Benchmark Scores

BenchmarkScoreSourceDate
ARC-C91.0%Benchgen evaluation2024-08
GSM8K88.7%Benchgen evaluation2024-08
MMLU78.5%Benchgen evaluation2024-08
MATH59.5%Benchgen evaluation2024-08
MMLU-Pro45.3%Benchgen evaluation2024-08
Arena Hard37.9%Benchgen evaluation2024-08

Phi-3.5 MoE vs Alternatives

ModelMMLUGSM8KMATHLicense
Phi-3.5 MoE Instruct78.5%88.7%59.5%Apache 2.0
Phi-4MIT
Phi-4 Reasoning97.2%92.7%Apache 2.0
Phi-3.5 Mini InstructApache 2.0

Phi-3.5 MoE vs Phi-4 Reasoning: Phi-4 Reasoning dominates on MATH (92.7% vs 59.5%) and GSM8K (97.2% vs 88.7%). Use Phi-3.5 MoE only for legacy compatibility; prefer Phi-4 Reasoning for new deployments.

Frequently Asked Questions

What is Phi-3.5 MoE Instruct? Microsoft's August 2024 MoE instruction model scoring 91.0% ARC-C, 88.7% GSM8K, 78.5% MMLU. Apache 2.0 — superseded by Phi-4 Reasoning for math tasks.

Specs from Microsoft's Phi-3.5 MoE Instruct release (August 2024) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.