Benchgen
Models/deepseek/

DeepSeek-R1-0528

DraftPublic

Model Details

DeepSeek R1 0528

Organization License Released Reasoning

Quick answer: DeepSeek R1 0528 is DeepSeek's May 28, 2025 update to R1, scoring 92.3% SimpleQA, 97.3% MATH, 96.2% GSM8K, 81.0% GPQA Diamond, 73.1% LiveCodeBench, and 85% MMLU-Pro under MIT license.

At a Glance

Where R1 0528 leads

  • 92.3% SimpleQA — exceptional factual accuracy
  • 97.3% MATH — near-perfect competition mathematics
  • 96.2% GSM8K — near-perfect grade-school math
  • 81.0% GPQA Diamond — strong graduate-level science reasoning
  • 73.1% LiveCodeBench — excellent real-world coding
  • MIT license — fully open including commercial use and modification

Where it lags

  • 58.0% Arena Hard v2 — moderate instruction quality vs frontier models
  • 35.1% BigCodeBench — lower on complex code generation
  • 5.7% TerminalBench — poor terminal task performance
  • Large model: high inference cost

Best for: Open-source reasoning pipelines requiring frontier math, STEM, and factual accuracy; research teams needing MIT-licensed reasoning.

What DeepSeek R1 0528 Is

DeepSeek R1 0528 is a checkpoint update of DeepSeek's R1 reasoning model, released May 28, 2025. The 0528 suffix indicates the release date. Like the original R1, it uses a chain-of-thought reasoning architecture via GRPO (Group Relative Policy Optimisation) training.

The 92.3% SimpleQA score is among the highest publicly reported — reflecting R1's factual grounding alongside reasoning. Combined with near-perfect MATH (97.3%) and GSM8K (96.2%), this checkpoint is one of the strongest open-weight reasoning models for mathematics.

The MIT license makes R1 0528 commercially unrestricted — teams can fine-tune, distil, and redistribute derivatives.

Specifications

FieldValue
OrganizationDeepSeek
LicenseMIT
HuggingFacedeepseek-ai/DeepSeek-R1-0528
Release dateMay 28, 2025
ArchitectureChain-of-thought (GRPO)
ModalityText only

Pricing

Open weights under MIT — self-host at no cost. Available via DeepSeek API and third-party providers (Together AI, Fireworks, etc.).

Public Benchmark Scores

BenchmarkScoreSourceDate
SimpleQA92.3%Benchgen evaluation2025-05
MATH97.3%Benchgen evaluation2025-05
GSM8K96.2%Benchgen evaluation2025-05
GPQA Diamond81.0%Benchgen evaluation2025-05
LiveCodeBench73.1%Benchgen evaluation2025-05
MMLU-Pro85%Benchgen evaluation2025-05
Arena Hard v258.0%Benchgen evaluation2025-05
AI2 Reasoning Challenge98.1%Benchgen evaluation2025-05

DeepSeek R1 0528 vs Alternatives

ModelSimpleQAMATHLiveCodeBenchLicense
DeepSeek R1 052892.3%97.3%73.1%MIT
DeepSeek-V3.2 Exp97.1%74.1%MIT
QwQ-32BApache 2.0
o3-mini97.9%Proprietary

R1 0528 vs DeepSeek-V3.2 Exp: near-equal LiveCodeBench (73.1% vs 74.1%). R1 0528 leads on explicit reasoning tasks; V3.2 Exp on instruction-following. Both MIT licensed.

Frequently Asked Questions

What is DeepSeek R1 0528? DeepSeek's May 2025 update to R1, scoring 92.3% SimpleQA, 97.3% MATH, 73.1% LiveCodeBench, and 81.0% GPQA Diamond. MIT licensed reasoning model.
Is R1 0528 better than the original DeepSeek R1? R1 0528 is a checkpoint update (May 28, 2025) of the original R1 with improved performance on several benchmarks. The MIT license and specific benchmark improvements make it the recommended version over the original R1.

Specs from DeepSeek's R1 0528 release (May 2025) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.