Model launches
Introducing HIPAA-compliant DeepSeek-V4.1-Flash
DeepSeek-V4.1-Flash pairs a Causal Encoder-Decoder architecture with 8B/16B asymmetric activation, native vision, and 1M context, available on our design-partner program.
TL;DR
DeepSeek-V4.1-Flash is now available under a signed BAA with zero data retention - a 552B multimodal MoE that leads the open-weight field on nearly every agentic benchmark in DeepSeek's comparison table, with a KV cache roughly a quarter the size of V4-Flash's, while still trailing the closed frontier on the hardest terminal suites.
Every clinical document workflow reads far more than it writes. A prior-authorization packet, a chart export, a scanned letter, twelve months of claims: the input dwarfs the output, and for high-volume document work the input side drives the bill. DeepSeek-V4.1-Flash is the first release we have seen that attacks that asymmetry in the architecture rather than through a discount. DeepSeek's own figure for it: a global KV cache of 890 bytes per token, about a quarter of what DeepSeek-V4-Flash needs at the same context depth.
Today, DeepSeek-V4.1-Flash is available on our HIPAA-compliant design-partner program under a signed BAA, with end-to-end encryption and strict zero data retention on prompts and outputs. It is the first release in DeepSeek's Causal Encoder-Decoder line, and the strongest open-weight agentic model we have hosted. See the DeepSeek-V4.1-Flash model page →
What is DeepSeek-V4.1-Flash?
DeepSeek published the weights on September 10, 2026, under the MIT
license (deepseek-ai/DeepSeek-V4.1-Flash). The backbone is a 552B
parameter Mixture-of-Experts model: a 40-layer transformer split into a
20-layer causal encoder and a 20-layer decoder, with 1 shared and 384
routed experts per layer, 6 active per token. It is multimodal - a
DeepSeek-ViT encoder (2D-RoPE, 3x3 pixel-unshuffle) trained jointly with
text from the start of pre-training, on a 45T-token corpus, not bolted
on afterward.
The Causal Encoder-Decoder layout is where the economics change. The decoder's global KV cache is projected from the final encoder hidden states instead of being derived at every decoder layer, which lets the model activate 8B parameters per token during prefill and 16B during decode. SWA Bounded Replay then reconstructs sliding-window KV states by replaying only the most recent window instead of persisting them to SSD, cutting the persistent KV footprint to roughly an eighth of V4-Flash's.
Compressed Sparse Attention 2 (CSA2) does the rest. Each attention layer takes one of three static modes - Full, Reindex, or Reuse - so layers share main KV and indexer tensors, and a Hierarchical Sparse Indexer keeps deeper indexing cost bounded as context grows. With FP4 main KV caching (E2M1, one E4M3 scale per 16 channels), DeepSeek measures the global cache at 890 bytes per token, roughly a quarter of V4-Flash's and about 437-fold below that of its first-generation model. The rest of the stack is Single-Pass mHC residual mixing, Engram conditional memory (196B sparsely-accessed parameters), and DSpark speculative decoding inside the checkpoint.
Two operational details matter for integration. Reasoning effort is
continuously controllable from 1 to 100 (an integer, not a three-level
switch); DeepSeek's published results all use reasoning_effort=100.
Recommended sampling is temperature 1.0, top_p 0.95 or 1.0, and
max_tokens 256K or higher against the 1M-token context window.
Performance profile: what it's good at vs. what it's not
What DeepSeek-V4.1-Flash excels at
- Agentic coding at Flash-class cost. Terminal-Bench 2.1 at 90.6
(DeepSeek Harness Minimal, 1M context,
reasoning_effort=100) tops the model card's own comparison table, which includes Opus-5.0 (89.1) and GPT-5.6 Sol (88.8). DeepSWE v1.1 at 74.2 and NL2Repo-Bench at 64.0 follow at the same effort setting. - Long-horizon automation. AutomationBench 54.8 beats Opus-5.0's 50.3 and GPT-5.6 Sol's 45.8, and Agent's Last Exam 31.8 beats every other model in the table. Claim assembly, referral routing, and back-office workflow agents are exactly this benchmark shape.
- Security-adjacent agentic depth. CyberGym 88.1 is the highest score in the card's table (GLM-5.3 and GPT-5.6 Sol both 84.5), and SEC-Bench Pro 62.8 is nearly seven points above V4-Pro's 56.4.
- Long-context document work. A 1M-token window on a KV cache of 890 bytes per token is the combination that makes whole-packet reading affordable: DocVQA 95.6 and MMMU-Pro 56.5 on DeepSeek's base-model table, from the same weights.
- Native vision in one checkpoint. Chartography with tools 78.9 and BabyVision with tools 89.6 come from the weights that posted the coding scores, under the Claude Code harness at 512K context. There is no separate vision checkpoint to add to a deployment map.
Limitations to keep in mind
- The hardest terminal suites are not close. Terminal-Bench 3.0 at 30.0 leads the other open models in the table, but 4.0 at 31.2 sits behind GLM-5.3's 37.9, and both rows trail Opus-5.0 (43.3 and 51.8) and GPT-5.6 Sol (34.4 and 39.9) by a wide margin. The 90.6 headline belongs to the 2.1 suite.
- Deep reasoning is a tier behind. HLE 36.8 (39.1 with tools) against Opus-5.0's 56.3, and GPQA Diamond 90.9 behind Opus-5.0, GPT-5.6 Sol, and K3. It reads as a stronger agent than it is a scientist.
- Verbose and output-heavy at max effort. Artificial Analysis
measured 250M generated tokens across its Intelligence Index, well
above the 140M median, at a cost of $0.27 per task on DeepSeek's own
API. At
reasoning_effort=100, spend lands on the output side. - Vision reads documents, not images for diagnosis. Chart and screenshot understanding does not make this a radiology or pathology tool, and we make no clinical accuracy claim for the vision path.
- Serving it needs the reference encoder. DeepSeek ships no
Jinja-format chat template; integration runs through the reference
encoding.pyor thedeepseek-recipetoolkit. Teams self-hosting should budget that work, and re-validate any forked DSpark module, which now lives inside the checkpoint.
The benchmarks: how DeepSeek-V4.1-Flash compares
All figures below are from DeepSeek's model card for V4.1-Flash
(vendor-reported; code-agent rows use DeepSeek Harness Minimal at
reasoning_effort=100 with a 1M-token window, DeepSWE v1.1 uses mini-SWE,
SEC-Bench Pro uses Claude Code, and the visual agent rows use Claude Code
at 512K). A dash means the benchmark was not published for that model.
| Benchmark (Focus) | V4.1-Flash | V4-Pro | V4-Flash | Opus-5.0 | GPT-5.6 Sol | K3 | GLM-5.3 |
|---|---|---|---|---|---|---|---|
| Terminal-Bench 2.1 (agentic coding) | 90.6 | 87.9 | 82.7 | 89.1 | 88.8 | 88.3 | 88.2 |
| Terminal-Bench 3.0 (harder terminal) | 30.0 | 11.8 | 7.6 | 43.3 | 34.4 | 17.7 | 28.3 |
| Terminal-Bench 4.0 (harder terminal) | 31.2 | 12.4 | 7.0 | 51.8 | 39.9 | 12.6 | 37.9 |
| DeepSWE v1.1 (software engineering) | 74.2 | 62.7 | 54.4 | 74.0 | 73.0 | 67.5 | 66.9 |
| NL2Repo-Bench (repo-level coding) | 64.0 | 61.5 | 54.2 | 75.3 | 56.8 | 58.0 | 58.0 |
| CyberGym (security reasoning) | 88.1 | 83.3 | 76.7 | - | 84.5 | 80.0 | 84.5 |
| SEC-Bench Pro (security agentic) | 62.8 | 56.4 | 30.9 | - | 74.3 | - | - |
| AutomationBench (automation) | 54.8 | 43.2 | 37.7 | 50.3 | 45.8 | 46.7 | 48.8 |
| Agent's Last Exam (agentic) | 31.8 | 25.7 | 25.2 | 28.6 | 26.7 | 27.6 | 28.5 |
| HLE w/ tools (multidisciplinary) | 63.9 | 60.0 | 51.5 | 63.6 | - | 59.8 | 62.5 |
| HLE (multidisciplinary) | 36.8 (39.1†) | 42.7† | 37.8† | 56.3 | 44.5 | 43.5 | 42.0† |
| GPQA Diamond (science reasoning) | 90.9 | 92.4 | 89.9 | 93.4 | 94.1 | 92.9 | 88.1 |
| Codeforces (competitive coding) | 3471 | 3348 | 3289 | - | - | - | - |
| Chartography w/ tools (chart reading) | 78.9 | - | - | 84.0 | 79.9 | 68.1 | - |
| BabyVision w/ tools (visual agent) | 89.6 | - | - | 94.1 | 88.9 | 85.7 | - |
† Text-only subset of HLE. Baseline columns are as published in DeepSeek's card; other vendors may report different figures for the same benchmark under their own harnesses and judges.
Two readings matter. On Terminal-Bench 3.0 and 4.0, V4.1-Flash improves sharply on its V4 predecessors (30.0 and 31.2 against 11.8 and 12.4 for V4-Pro) and still trails Opus-5.0 by 13 and 21 points - that is the real gap, and the 2.1-suite headline does not carry across. On GPQA and HLE, it sits roughly a tier below the closed frontier. Put plainly: V4.1-Flash leads the open-weight field on nearly every agentic row in this table, and it does not lead the frontier overall.
Artificial Analysis, which measures independently, scores it 39 on the Intelligence Index against a median of 19 for open-weight models of comparable size, and measures 206 tokens per second of output - fast for this class, and verbose.
Why run DeepSeek-V4.1-Flash via our HIPAA-compliant API?
- Enterprise-grade privacy. Available under our HIPAA-compliant design-partner program with a signed BAA, end-to-end data encryption, and strict zero-data-retention on prompts and outputs. Scanned packets, chart exports, and claims data are exactly the inputs this model is built to read, and none of that data is used for training. General availability is coming soon.
- Uncompromised performance. We run the serving stack this architecture needs: asymmetric prefill/decode kernels, FP4 KV caching, SWA Bounded Replay, DSpark speculation, and 1M-token context scaling. You get the cache economics without operating the cluster that produces them.
- Transparent pricing. The KV footprint is the cost driver for input-heavy pipelines, and at 890 bytes per token this generation moves it well below every prior DeepSeek checkpoint we host. Our serving layer bills $0.30 per 1M input tokens, $0.006 per 1M cached input tokens, and $1.20 per 1M output tokens; the cached-input rate is a 98% discount on input, which is the number that matters for repeated document context.
Ready to put long-context agentic work under your BAA? Join the waitlist now.
Run it HIPAA-compliant
DeepSeek-V4.1-Flash on OpenMed Router
$0.30/M input · $1.20/M output · 1048576 context
Chris Williams, MD
Chris Williams, MD is a physician, clinical AI researcher and the co-founder of OpenMed Router, working to make open source AI models safely accessible to healthcare organizations under HIPAA. He writes about clinical AI, model selection, compliance, and the practical adoption of open source inference in clinical and operational workflows.
Join the waitlist