Benchgen
Models/deepseek/

DeepSeek-V4.1-Flash

DraftPublic

Model Details

DeepSeek-V4.1-Flash

Org Pricing License Modality Released

Quick answer: DeepSeek-V4.1-Flash is an open-weight (MIT-licensed) multimodal MoE model using a novel Causal Encoder-Decoder (CED) architecture — a 552B-parameter backbone that activates only 8B params during prefill and 16B during decode, plus a 196B-parameter sparsely-accessed Engram memory module. It supports a 1M-token context and continuously controllable reasoning effort (1–100).

At a Glance

  • Where it leads: HumanEval (79.4% Pass@1), DeepSWE v1.1 (74.2% resolved), Terminal-Bench 2.1 (90.6%), CyberGym (88.1%), DocVQA (95.6%) among the models compared in its own technical report.
  • Where it lags: Terminal-Bench 3.0 (30.0%) and Terminal-Bench 4.0 (31.2%) trail several closed frontier models on the same eval.

What DeepSeek-V4.1-Flash Is

DeepSeek-V4.1-Flash pairs a 40-layer Transformer (20-layer causal encoder + 20-layer decoder) with 1 shared expert and 384 routed experts per MoE layer (6 activated per token), SWA Bounded Replay, Compressed Sparse Attention 2 (CSA2), FP4 KV caching, Single-Pass mHC, and DSpark speculative decoding. Trained on 45T tokens (context extended from 64K to 1M at the 34T-token mark).

Specifications

OrganizationDeepSeek
LicenseMIT (open weights)
Release dateSeptember 10, 2026 (HF); technical report Sep 18, 2026
ModalityMultimodal (text + vision)
Knowledge cutoffNot disclosed

Public Benchmark Scores

BenchmarkScoreSourceDate
MMLU-Pro74.1 (5-shot, base)DeepSeekSep 2026
C-Eval92.1 (5-shot, base)DeepSeekSep 2026
SuperGPQA53.1 (5-shot, base)DeepSeekSep 2026
Big-Bench Hard86.1 (3-shot, base)DeepSeekSep 2026
DROP87.9 F1 (1-shot, base)DeepSeekSep 2026
HellaSwag87.2 (0-shot, base)DeepSeekSep 2026
BigCodeBench60.6 Pass@1 (3-shot, base)DeepSeekSep 2026
HumanEval79.4 Pass@1 (0-shot, base)DeepSeekSep 2026
GSM8K93.0 (8-shot, base)DeepSeekSep 2026
MATH61.1 (4-shot, base)DeepSeekSep 2026
MGSM80.2 (8-shot, base)DeepSeekSep 2026
LongBench-V245.2 (1-shot, base)DeepSeekSep 2026
MMMU-Pro56.5 (4-shot, base)DeepSeekSep 2026
DocVQA95.6 LLM-Judge (4-shot, base)DeepSeekSep 2026
GPQA Diamond90.9 Pass@1 (max effort)DeepSeekSep 2026
Humanity's Last Exam63.9 Pass@1 (with tools, max effort)DeepSeekSep 2026
MathArena Apex65.6 Pass@1 (max effort)DeepSeekSep 2026
Terminal-Bench 2.190.6 Pass@1 (max effort)DeepSeekSep 2026
Terminal-Bench 3.030.0 Pass@1 (max effort)DeepSeekSep 2026
Terminal-Bench 4.031.2 Pass@1 (max effort)DeepSeekSep 2026
DeepSWE v1.174.2 Resolved (max effort)DeepSeekSep 2026
ProgramBench20.3 Almost@1 (max effort)DeepSeekSep 2026
NL2Repo-Bench64.0 (max effort)DeepSeekSep 2026
CyberGym88.1 Pass@1 (max effort)DeepSeekSep 2026
SEC-Bench Pro62.8 Pass@1 (max effort)DeepSeekSep 2026
ExploitGym15.3 Pass@1 (max effort)DeepSeekSep 2026
Agents' Last Exam31.8 Pass@1 (max effort)DeepSeekSep 2026
BabyVision89.6 Pass@1 (with tools, max effort)DeepSeekSep 2026
ZeroBench49.0 Pass@5 (main, with tools, max effort)DeepSeekSep 2026

Scores as reported in DeepSeek's own DeepSeek-V4.1-Flash model card and technical report (huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash, September 2026); not independently verified by Benchgen.