The V4 family's first image-input model: an experimental, API-only multimodal build of DeepSeek-V4-Flash-0731 that accepts mixed text + image input (JPEG, PNG, GIF, WebP; up to 600 images per request). DeepSeek says it "matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge" and makes "a major leap over V4-Flash" on multimodal agent benchmarks, "bringing multimodal agent performance close to Opus-4.8" (no scores published). Served as deepseek-v4-flash-vision-exp with a 1M-token context, 384K max output, thinking mode on by default with low/high/max reasoning_effort, and Chat Completions, Anthropic Messages and Responses API support. Weights followed on HuggingFace on September 1, 2026 under the MIT license (~305B by safetensors count — the 284B V4-Flash MoE backbone plus the vision tower; deepseek_v4, 43 layers, 256 routed experts, top-6, 1,048,576 positions).

Images are auto-resized to roughly 800×800 pixels and tokenized at up to 384 tokens per image, billed at V4-Flash text rates ($0.22/M input, $0.66/M output off-peak). Launched alongside a free Files API (upload an image once, reference it by file_id) and DeepSeek Harness 0.1.1, which added the model to its DeepSeek adapter the same day. Artificial Analysis lists it as a proprietary model and scores its max-effort reasoning mode within a point of the text-only V4-Flash-0731 (AA Intelligence Index v4.3: 35 vs 41).

Model Details

Context window 1,048,576
AA Intelligence 35 was 42 on v4.2
License MIT
Base model deepseek-v4-flash
multimodalvisionreasoningagenticproprietary

Related