Quick answer: DeepSeek R1 0528 is DeepSeek's May 28, 2025 update to R1, scoring 92.3% SimpleQA, 97.3% MATH, 96.2% GSM8K, 81.0% GPQA Diamond, 73.1% LiveCodeBench, and 85% MMLU-Pro under MIT license.
Where R1 0528 leads
Where it lags
Best for: Open-source reasoning pipelines requiring frontier math, STEM, and factual accuracy; research teams needing MIT-licensed reasoning.
DeepSeek R1 0528 is a checkpoint update of DeepSeek's R1 reasoning model, released May 28, 2025. The 0528 suffix indicates the release date. Like the original R1, it uses a chain-of-thought reasoning architecture via GRPO (Group Relative Policy Optimisation) training.
The 92.3% SimpleQA score is among the highest publicly reported — reflecting R1's factual grounding alongside reasoning. Combined with near-perfect MATH (97.3%) and GSM8K (96.2%), this checkpoint is one of the strongest open-weight reasoning models for mathematics.
The MIT license makes R1 0528 commercially unrestricted — teams can fine-tune, distil, and redistribute derivatives.
| Field | Value |
|---|---|
| Organization | DeepSeek |
| License | MIT |
| HuggingFace | deepseek-ai/DeepSeek-R1-0528 |
| Release date | May 28, 2025 |
| Architecture | Chain-of-thought (GRPO) |
| Modality | Text only |
Open weights under MIT — self-host at no cost. Available via DeepSeek API and third-party providers (Together AI, Fireworks, etc.).
| Benchmark | Score | Source | Date |
|---|---|---|---|
| SimpleQA | 92.3% | Benchgen evaluation | 2025-05 |
| MATH | 97.3% | Benchgen evaluation | 2025-05 |
| GSM8K | 96.2% | Benchgen evaluation | 2025-05 |
| GPQA Diamond | 81.0% | Benchgen evaluation | 2025-05 |
| LiveCodeBench | 73.1% | Benchgen evaluation | 2025-05 |
| MMLU-Pro | 85% | Benchgen evaluation | 2025-05 |
| Arena Hard v2 | 58.0% | Benchgen evaluation | 2025-05 |
| AI2 Reasoning Challenge | 98.1% | Benchgen evaluation | 2025-05 |
| Model | SimpleQA | MATH | LiveCodeBench | License |
|---|---|---|---|---|
| DeepSeek R1 0528 | 92.3% | 97.3% | 73.1% | MIT |
| DeepSeek-V3.2 Exp | 97.1% | — | 74.1% | MIT |
| QwQ-32B | — | — | — | Apache 2.0 |
| o3-mini | — | 97.9% | — | Proprietary |
R1 0528 vs DeepSeek-V3.2 Exp: near-equal LiveCodeBench (73.1% vs 74.1%). R1 0528 leads on explicit reasoning tasks; V3.2 Exp on instruction-following. Both MIT licensed.
Specs from DeepSeek's R1 0528 release (May 2025) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.