Quick answer: DeepSeek-V4.1-Flash is an open-weight (MIT-licensed) multimodal MoE model using a novel Causal Encoder-Decoder (CED) architecture — a 552B-parameter backbone that activates only 8B params during prefill and 16B during decode, plus a 196B-parameter sparsely-accessed Engram memory module. It supports a 1M-token context and continuously controllable reasoning effort (1–100).
DeepSeek-V4.1-Flash pairs a 40-layer Transformer (20-layer causal encoder + 20-layer decoder) with 1 shared expert and 384 routed experts per MoE layer (6 activated per token), SWA Bounded Replay, Compressed Sparse Attention 2 (CSA2), FP4 KV caching, Single-Pass mHC, and DSpark speculative decoding. Trained on 45T tokens (context extended from 64K to 1M at the 34T-token mark).
| Organization | DeepSeek |
| License | MIT (open weights) |
| Release date | September 10, 2026 (HF); technical report Sep 18, 2026 |
| Modality | Multimodal (text + vision) |
| Knowledge cutoff | Not disclosed |
| Benchmark | Score | Source | Date |
|---|---|---|---|
| MMLU-Pro | 74.1 (5-shot, base) | DeepSeek | Sep 2026 |
| C-Eval | 92.1 (5-shot, base) | DeepSeek | Sep 2026 |
| SuperGPQA | 53.1 (5-shot, base) | DeepSeek | Sep 2026 |
| Big-Bench Hard | 86.1 (3-shot, base) | DeepSeek | Sep 2026 |
| DROP | 87.9 F1 (1-shot, base) | DeepSeek | Sep 2026 |
| HellaSwag | 87.2 (0-shot, base) | DeepSeek | Sep 2026 |
| BigCodeBench | 60.6 Pass@1 (3-shot, base) | DeepSeek | Sep 2026 |
| HumanEval | 79.4 Pass@1 (0-shot, base) | DeepSeek | Sep 2026 |
| GSM8K | 93.0 (8-shot, base) | DeepSeek | Sep 2026 |
| MATH | 61.1 (4-shot, base) | DeepSeek | Sep 2026 |
| MGSM | 80.2 (8-shot, base) | DeepSeek | Sep 2026 |
| LongBench-V2 | 45.2 (1-shot, base) | DeepSeek | Sep 2026 |
| MMMU-Pro | 56.5 (4-shot, base) | DeepSeek | Sep 2026 |
| DocVQA | 95.6 LLM-Judge (4-shot, base) | DeepSeek | Sep 2026 |
| GPQA Diamond | 90.9 Pass@1 (max effort) | DeepSeek | Sep 2026 |
| Humanity's Last Exam | 63.9 Pass@1 (with tools, max effort) | DeepSeek | Sep 2026 |
| MathArena Apex | 65.6 Pass@1 (max effort) | DeepSeek | Sep 2026 |
| Terminal-Bench 2.1 | 90.6 Pass@1 (max effort) | DeepSeek | Sep 2026 |
| Terminal-Bench 3.0 | 30.0 Pass@1 (max effort) | DeepSeek | Sep 2026 |
| Terminal-Bench 4.0 | 31.2 Pass@1 (max effort) | DeepSeek | Sep 2026 |
| DeepSWE v1.1 | 74.2 Resolved (max effort) | DeepSeek | Sep 2026 |
| ProgramBench | 20.3 Almost@1 (max effort) | DeepSeek | Sep 2026 |
| NL2Repo-Bench | 64.0 (max effort) | DeepSeek | Sep 2026 |
| CyberGym | 88.1 Pass@1 (max effort) | DeepSeek | Sep 2026 |
| SEC-Bench Pro | 62.8 Pass@1 (max effort) | DeepSeek | Sep 2026 |
| ExploitGym | 15.3 Pass@1 (max effort) | DeepSeek | Sep 2026 |
| Agents' Last Exam | 31.8 Pass@1 (max effort) | DeepSeek | Sep 2026 |
| BabyVision | 89.6 Pass@1 (with tools, max effort) | DeepSeek | Sep 2026 |
| ZeroBench | 49.0 Pass@5 (main, with tools, max effort) | DeepSeek | Sep 2026 |
Scores as reported in DeepSeek's own DeepSeek-V4.1-Flash model card and technical report (huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash, September 2026); not independently verified by Benchgen.