← Back to blog

Model launches

Introducing HIPAA-compliant DeepSeek-V4-Flash-Vision-Exp

DeepSeek's first multimodal model, Flash-Vision-Exp, keeps the Flash speed profile and adds image understanding under a signed BAA with zero data retention.

Chris Williams, MD

TL;DR

DeepSeek-V4-Flash-Vision-Exp, DeepSeek's first multimodal model, is now available under a signed BAA with zero data retention - Flash-class speed and pricing plus native image understanding, for workloads that start from a scan, a chart, or a photo.

Healthcare workloads rarely start from a clean text prompt. They start from a scanned insurance card, a chart in an EHR screenshot, a photograph of a wound, a graph in a lab report, or a scanned prior-auth letter. Every text-only model we have covered on this blog - including every DeepSeek release to date - has the same hard boundary: it cannot see.

Today, we are expanding what our design-partner program can see. DeepSeek-V4-Flash-Vision-Exp, the first multimodal model in the DeepSeek-V4 family, is now available on our HIPAA-compliant API platform under a signed BAA, with end-to-end encryption and strict zero data retention - the same as every model we host.

What is DeepSeek-V4-Flash-Vision-Exp?

DeepSeek published the weights on August 31, 2026, under the MIT license, ten days after an API-only rollout (the weights surfaced publicly around August 21). The model keeps the V4-Flash architecture - a 284B sparse MoE, roughly 13B parameters active per token, 1M token context window, hybrid attention (CSA + HCA) with KV cache near 10% of a standard footprint - and adds visual modules with continued training to unlock image understanding. The hosted checkpoint registers at 305B parameters including the DSpark speculative decoding module, matching what the DeepSeek card describes.

Two design points matter for how we think about it clinically. First, the vision encoder is integrated into the existing Flash architecture rather than bolted on, so serving cost stays close to the text-only model. Second, this is clearly labelled experimental - Exp is in the name, and DeepSeek frames it as their first multimodal release, not a finished frontier vision model. That framing is accurate, and we treat it that way below.

Performance profile: what it's good at vs. what it's not

What Flash-Vision-Exp excels at

  • Vision-aware agents, not just vision chat. On Chartography with tools it scores 64.3 against Opus-4.8's 65.0 - statistical chart-reading, which is what a lab trend graph or a monthly utilization dashboard actually is. On ZeroBench with tools (Pass@5) it scores 35.0, ahead of Opus-4.8's 34.0.
  • Text agentic capability that barely regressed. Terminal-Bench 2.1 moved from 82.7 to 83.9, NL2Repo from 54.2 to 57.7, DeepSWE from 54.4 to 59.3 (all DeepSeek-reported, DeepSeek Harness minimal mode at max effort). Adding vision usually costs coding performance; here it largely did not.
  • Multimodal tool workflows. On Agents' Last Exam it reaches 27.3, past Opus-4.8's 25.7, with inputs it actually ingests images into (the Flash-0731 baseline "solves" the same benchmark by ignoring them).
  • Cost profile. It keeps the Flash-class serving economics of the text-only checkpoint, which is what makes vision at this speed interesting for high-volume document work rather than a premium tier.
  • A first for the family. For teams standardized on the V4-Flash API shape, this is a drop-in way to add image input without re-platforming onto a different vendor's model.

Limitations to keep in mind

  • Experimental by name, and by behavior. On multimodal tasks where a score is known, it lands near but rarely past Opus-4.8 - except on ZeroBench, where pass@5 hides high per-run variance that is worth treating as noise until third-party reruns land.
  • Not for diagnostic imaging. Despite the healthcare framing, nothing about this model is a radiology or pathology tool, and we make no claim that it substitutes for one. It reads charts, documents, and screenshots - not X-rays, CTs, or slides with clinical-grade reliability.
  • Text-heavy agentic work still prefers the Pro checkpoint. At 83.9 on Terminal-Bench, this remains a Flash-tier model. If your work is pure code or pure long-doc reasoning and not image-touching, our DeepSeek-V4-Flash and DeepSeek-V4-Pro options remain better value for dollar per token.
  • Latency under vision load. Feeding images at Flash-class cost implies Flash-class throughput - fine for background pipelines, but interactive chat sessions mixing large images will feel slower than text-only sessions, and we have not independently benchmarked that here (marking this as our own observation, not a measured score).
  • Short support window. An -Exp model on a fast-moving family can be superseded quickly. If you build on it, keep a rollback path to text-only Flash in your pipeline plan.

The benchmarks: how Flash-Vision-Exp compares

Benchmarks below are from the official DeepSeek model card on Hugging Face (vendor-reported, agentic text tasks evaluated with DeepSeek Harness minimal mode at max reasoning effort); the - entries mean the benchmark was not run on that checkpoint.

Benchmark (Focus)Flash-Vision-ExpFlash-0731Opus-4.8
Terminal-Bench 2.1 (agentic coding)83.982.785.0
NL2Repo (repo-level coding)57.754.269.7
DeepSWE (software engineering)59.354.458.0
Cybergym (security reasoning)75.376.778.3
Toolathlon-Verified (tool use)75.970.376.2
DSBench-Hard (data-adjacent)63.659.671.7
ApexBench Pass@1 (multimodal agent)36.526.2†39.4
Agents' Last Exam (multimodal agent)27.325.2†25.7
Chartography (chart understanding)64.3-65.0
ZeroBench Pass@5 (hard visual)35.0-34.0

† The Flash-0731 baseline ignores multimodal elements in the input, per DeepSeek's footnote - treat those rows as apples-to-oranges for the vision workloads specifically.

Why run Flash-Vision-Exp via our HIPAA-compliant API?

  • Enterprise-grade privacy. Available under our HIPAA-compliant design-partner program with a signed BAA, end-to-end data encryption, and strict zero-data-retention on prompts and outputs - including the images. Scanned documents, screenshots, and photos of clinical paperwork are exactly the data you cannot paste into a consumer AI tool, and exactly what this model is designed to read. Your patient data is never used for training. General availability is coming soon.
  • Uncompromised performance. We handle the vision-aware serving stack, the DSpark speculative decoding, and 1M-token context-window scaling - you get Flash-class multimodal agents without operating a GPU cluster.
  • Transparent pricing. Flash-Vision-Exp keeps the Flash-class economics of the text-only checkpoint at our serving layer ($0.14 per 1M input / $0.28 per 1M output on DeepSeek's MIT-licensed family) - a fraction of closed multimodal pricing on Azure OpenAI or AWS Bedrock.

Ready to give your document pipelines actual eyes under your BAA? Join the waitlist now.

Run it HIPAA-compliant

DeepSeek-V4-Flash-Vision-Exp on OpenMed Router

$0.14/M input · $0.28/M output · 1048576 context

Chris Williams, MD

Chris Williams, MD is a physician, clinical AI researcher and the co-founder of OpenMed Router, working to make open source AI models safely accessible to healthcare organizations under HIPAA. He writes about clinical AI, model selection, compliance, and the practical adoption of open source inference in clinical and operational workflows.

Join the waitlist

Be first in line for HIPAA-compliant open source inference