Benchgen
Models/deepseek/

DeepSeek-V4 Flash Vision Exp

DraftPublic

Model Details

DeepSeek-V4 Flash Vision Exp

Organization Modality License Released

Quick answer: DeepSeek-V4-Flash-Vision-Exp is DeepSeek's first experimental multimodal model in the V4 family, built on the DeepSeek-V4-Flash architecture with added visual encoder/aligner modules and continued training. It improves multimodal agent capability substantially over the text-only DeepSeek-V4-Flash-0731 baseline while holding text-only agent performance roughly steady. Released under the MIT license, September 2026.

At a Glance

Where it leads

  • Clear multimodal-agent gains over its own text-only predecessor: ApexBench Pass@1 jumps to 36.5% (from 26.2% when the predecessor ignores image input), Agents' Last Exam to 27.3% (from 25.2%)
  • Text-only agent capability held essentially flat vs. DeepSeek-V4-Flash-0731 (e.g. Terminal Bench 2.1: 83.9% vs 82.7%), so vision was added without a text-capability tax
  • ZeroBench (Pass@5) at 35.0% edges out Claude Opus 4.8's 34.0% — notable for an experimental release

Where it lags

  • Still trails Claude Opus 4.8 on most text-agent benchmarks (NL2Repo 57.7% vs 69.7%, DSBench-Hard 63.6% vs 71.7%)
  • Explicitly labeled "Exp" (experimental) by DeepSeek — not positioned as a production-grade flagship

Best for: teams evaluating early multimodal agent capability from DeepSeek's V4 line before a non-experimental release ships.

What DeepSeek-V4 Flash Vision Exp Is

DeepSeek-V4-Flash-Vision-Exp is DeepSeek's first attempt at bringing visual understanding into its V4-Flash architecture. Rather than a from-scratch multimodal model, it starts from the existing DeepSeek-V4-Flash-0731 checkpoint, adds a vision encoder and aligner, and continues training to unlock image understanding — a lower-cost path to multimodality that trades off against the ceiling a from-scratch multimodal design might reach.

The released repository ships a minimal reference PyTorch implementation covering the vision encoder/aligner, "DFlash" attention, MoE routing, Hyper-Connections, and the "DSpark" speculative-decoding forward path — DeepSeek publishes this alongside the weights specifically to let others reproduce inference rather than relying solely on hosted API access.

Specifications

FieldValue
OrganizationDeepSeek
ParametersUndisclosed (built on DeepSeek-V4-Flash architecture)
ArchitectureDeepSeek-V4-Flash + vision encoder/aligner, DFlash attention, MoE, Hyper-Connections, DSpark speculative decoding
LicenseMIT
Release date2026-09-01 (approximate, per HF "Updated" timestamp)
ModalityMultimodal (text + vision)

Pricing

Open weights — free to download and self-host under the MIT license; cost is inference/hosting only. No hosted API pricing published at release.

Public Benchmark Scores

Scores below are self-reported by DeepSeek in the model's own HF model card, evaluated against DeepSeek-V4-Flash-0731 (its own text-only predecessor) and Claude Opus 4.8. They are not Benchgen measurements. DeepSeek evaluates using the minimal mode of its own DeepSeek Harness, max reasoning effort, temperature=1.0, top_p=0.95.

Not mapped (excluded, no matching Benchgen benchmark entity): DSBench-Hard, ApexBench (Pass@1 multimodal variant), Chartography. These are novel/uncommon enough that no local benchmark config exists yet; excluded rather than force-mapped to a similarly-named entity. Given this is DeepSeek's first experimental multimodal release (not a flagship), new benchmark configs were not stood up for these three — revisit if a non-experimental successor or other labs start reporting the same benchmarks.

DeepSeek-V4 Flash Vision Exp vs Alternatives

ModelTerminal-Bench 2.1DeepSWEToolathlon Verified
DeepSeek-V4 Flash Vision Exp83.9%59.3%75.9%
DeepSeek-V4 Flash Max———
Claude Opus 4.885.0%58.0%76.2%

Rival scores are as reported in DeepSeek's own comparison table (same source as this model's scores above), not independently re-verified by Benchgen.

DeepSeek-V4-Flash-Vision-Exp lands close to Claude Opus 4.8 on several text-agent benchmarks despite being an experimental add-on to an existing text model — the interesting signal here isn't raw capability but that DeepSeek achieved it without a meaningful text-capability regression from adding vision.

How DeepSeek-V4 Flash Vision Exp Performs on Real Agent Tasks

The clearest agent signal is the multimodal-vs-ignoring-images comparison: DeepSeek-V4-Flash-0731 (the text-only predecessor) is evaluated on the same multimodal benchmarks by simply ignoring image inputs, and DeepSeek-V4-Flash-Vision-Exp beats it meaningfully (ApexBench 36.5% vs 26.2%, Agents' Last Exam 27.3% vs 25.2%). That's a genuine capability delta, not just a benchmark artifact — but builders evaluating this for production agent work should weigh the "Exp" label: DeepSeek has not positioned this as a finished, stable release, and the underlying parameter count and training details remain largely undisclosed.

Frequently Asked Questions

What is DeepSeek-V4 Flash Vision Exp? It's DeepSeek's first experimental multimodal model in the V4 family — the existing DeepSeek-V4-Flash text architecture with added vision encoder/aligner modules and continued training.
How much does DeepSeek-V4 Flash Vision Exp cost? It's open weights under the MIT license — free to download and self-host. No hosted API pricing has been published.
Is DeepSeek-V4 Flash Vision Exp open source? Yes, released under the MIT license, with a reference PyTorch inference implementation published alongside the weights.
Does adding vision hurt DeepSeek-V4 Flash's text performance? No — DeepSeek reports text-only agent benchmarks holding roughly flat vs. the text-only DeepSeek-V4-Flash-0731 predecessor (e.g. 83.9% vs 82.7% on Terminal-Bench 2.1).

Specs and scores sourced from DeepSeek's official model card. Third-party comparison scores attributed inline. Last updated 2026-09-08.