Quick answer: DeepSeek-V3.2 Experimental is DeepSeek's September 2025 updated MoE model, scoring 85% MMLU-Pro, 97.1% SimpleQA, 74.1% LiveCodeBench, and 37.7% TerminalBench. Under MIT at $0.27/$1.10 per 1M tokens, it represents a significant step up in coding and factual accuracy from the baseline V3.
Where DeepSeek-V3.2 Exp leads
Where it lags
Best for: Production coding tasks; factual Q&A requiring high accuracy; cost-efficient frontier performance; teams using DeepSeek-V3 who want improved coding/factual scores.
DeepSeek-V3.2 Experimental is an iterative update to the V3 MoE architecture released September 2025. The "Experimental" tag indicates a preview release before the stable V3.2 GA version.
The model's 97.1% SimpleQA score is exceptional — representing among the highest factual accuracy scores on that benchmark. This is paired with a 74.1% LiveCodeBench score, making it one of the strongest open models for coding at V3 pricing.
At $0.27/$1.10 per 1M tokens (same as V3), it offers frontier performance improvements at no additional cost increase for API users. MIT license and open weights maintain DeepSeek's commitment to open access.
| Field | Value |
|---|---|
| Organization | DeepSeek |
| Context window | 128,000 tokens |
| License | MIT |
| HuggingFace | deepseek-ai/DeepSeek-V3-2-Experimental |
| Release date | September 2025 (experimental) |
| Knowledge cutoff | June 2025 |
| Modality | Text only |
| Tier | Price |
|---|---|
| Input | $0.27 / 1M tokens |
| Output | $1.10 / 1M tokens |
| Benchmark | Score | Source | Date |
|---|---|---|---|
| MMLU-Pro | 85% | Benchgen evaluation | 2025-09 |
| SimpleQA | 97.1% | Benchgen evaluation | 2025-09 |
| LiveCodeBench | 74.1% | Benchgen evaluation | 2025-09 |
| TerminalBench | 37.7% | Benchgen evaluation | 2025-09 |
| Model | MMLU-Pro | LiveCodeBench | SimpleQA | Price (in/out) |
|---|---|---|---|---|
| DeepSeek-V3.2 Experimental | 85% | 74.1% | 97.1% | $0.27/$1.10 |
| DeepSeek-V3 | — | 27.2% | 24.9% | $0.27/$1.10 |
| GPT-4.1 | — | — | — | $2/$8 |
V3.2 Experimental vs V3: +63pp LiveCodeBench improvement (74.1% vs 27.2% — V3 score likely from earlier version), near-perfect SimpleQA (97.1% vs 24.9%), all at identical pricing. A substantial upgrade for coding tasks.
Specs from DeepSeek's official V3.2 Experimental release (September 2025) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.