Quick answer: DeepSeek-V4-Flash-Vision-Exp is DeepSeek's first experimental multimodal model in the V4 family, built on the DeepSeek-V4-Flash architecture with added visual encoder/aligner modules and continued training. It improves multimodal agent capability substantially over the text-only DeepSeek-V4-Flash-0731 baseline while holding text-only agent performance roughly steady. Released under the MIT license, September 2026.
Where it leads
Where it lags
Best for: teams evaluating early multimodal agent capability from DeepSeek's V4 line before a non-experimental release ships.
DeepSeek-V4-Flash-Vision-Exp is DeepSeek's first attempt at bringing visual understanding into its V4-Flash architecture. Rather than a from-scratch multimodal model, it starts from the existing DeepSeek-V4-Flash-0731 checkpoint, adds a vision encoder and aligner, and continues training to unlock image understanding — a lower-cost path to multimodality that trades off against the ceiling a from-scratch multimodal design might reach.
The released repository ships a minimal reference PyTorch implementation covering the vision encoder/aligner, "DFlash" attention, MoE routing, Hyper-Connections, and the "DSpark" speculative-decoding forward path — DeepSeek publishes this alongside the weights specifically to let others reproduce inference rather than relying solely on hosted API access.
| Field | Value |
|---|---|
| Organization | DeepSeek |
| Parameters | Undisclosed (built on DeepSeek-V4-Flash architecture) |
| Architecture | DeepSeek-V4-Flash + vision encoder/aligner, DFlash attention, MoE, Hyper-Connections, DSpark speculative decoding |
| License | MIT |
| Release date | 2026-09-01 (approximate, per HF "Updated" timestamp) |
| Modality | Multimodal (text + vision) |
Open weights — free to download and self-host under the MIT license; cost is inference/hosting only. No hosted API pricing published at release.
Scores below are self-reported by DeepSeek in the model's own HF model card, evaluated against DeepSeek-V4-Flash-0731 (its own text-only predecessor) and Claude Opus 4.8. They are not Benchgen measurements. DeepSeek evaluates using the minimal mode of its own DeepSeek Harness, max reasoning effort, temperature=1.0, top_p=0.95.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| Terminal-Bench 2.1 | 83.9 | DeepSeek model card | 2026-09 |
| NL2Repo | 57.7 | DeepSeek model card | 2026-09 |
| CyberGym | 75.3 | DeepSeek model card | 2026-09 |
| DeepSWE | 59.3 | DeepSeek model card | 2026-09 |
| Toolathlon Verified | 75.9 | DeepSeek model card | 2026-09 |
| AutomationBench (Public) | 25.7 | DeepSeek model card | 2026-09 |
| Agents' Last Exam | 27.3 | DeepSeek model card | 2026-09 |
| ZeroBench (Pass@5) | 35.0 | DeepSeek model card | 2026-09 |
Not mapped (excluded, no matching Benchgen benchmark entity): DSBench-Hard, ApexBench (Pass@1 multimodal variant), Chartography. These are novel/uncommon enough that no local benchmark config exists yet; excluded rather than force-mapped to a similarly-named entity. Given this is DeepSeek's first experimental multimodal release (not a flagship), new benchmark configs were not stood up for these three — revisit if a non-experimental successor or other labs start reporting the same benchmarks.
| Model | Terminal-Bench 2.1 | DeepSWE | Toolathlon Verified |
|---|---|---|---|
| DeepSeek-V4 Flash Vision Exp | 83.9% | 59.3% | 75.9% |
| DeepSeek-V4 Flash Max | — | — | — |
| Claude Opus 4.8 | 85.0% | 58.0% | 76.2% |
Rival scores are as reported in DeepSeek's own comparison table (same source as this model's scores above), not independently re-verified by Benchgen.
DeepSeek-V4-Flash-Vision-Exp lands close to Claude Opus 4.8 on several text-agent benchmarks despite being an experimental add-on to an existing text model — the interesting signal here isn't raw capability but that DeepSeek achieved it without a meaningful text-capability regression from adding vision.
The clearest agent signal is the multimodal-vs-ignoring-images comparison: DeepSeek-V4-Flash-0731 (the text-only predecessor) is evaluated on the same multimodal benchmarks by simply ignoring image inputs, and DeepSeek-V4-Flash-Vision-Exp beats it meaningfully (ApexBench 36.5% vs 26.2%, Agents' Last Exam 27.3% vs 25.2%). That's a genuine capability delta, not just a benchmark artifact — but builders evaluating this for production agent work should weigh the "Exp" label: DeepSeek has not positioned this as a finished, stable release, and the underlying parameter count and training details remain largely undisclosed.
Specs and scores sourced from DeepSeek's official model card. Third-party comparison scores attributed inline. Last updated 2026-09-08.