DeepSeek
DeepSeek-V4.1-Flash
Specs
- Input price
- $0.30/M
- Cached input price
- $0.006/M
- Output price
- $1.20/M
- Context window
- 1M tokens
DeepSeek's Causal Encoder-Decoder debut - a 552B multimodal MoE with 8B/16B asymmetric activation, native vision, and a KV cache roughly a quarter of V4-Flash's.
Capabilities
- 552B MoE (1 shared + 384 routed experts, 6 active per token) split into a 20-layer causal encoder and 20-layer decoder
- Terminal-Bench 2.1 at 90.6 (DeepSeek Harness Minimal, reasoning_effort=100) - top of the model card's comparison table
- AutomationBench 54.8 and Agent's Last Exam 31.8 lead the table; CyberGym 88.1 is the highest security score
- 890 bytes per token global KV cache (FP4 main KV, SWA Bounded Replay) at a 1M-token context window
- Native vision from pre-training: DocVQA 95.6, Chartography with tools 78.9, BabyVision with tools 89.6
Best for
- Input-heavy clinical document pipelines: prior-auth packets, chart exports, and claims history
- Long-horizon automation agents: claim assembly, referral routing, and back-office workflows
- Agentic coding and security-adjacent review at Flash-class cost
Limitations to keep in mind
- Trails the closed frontier on the hardest suites: Terminal-Bench 3.0 at 30.0 and 4.0 at 31.2
- Deep reasoning is a tier behind: GPQA Diamond 90.9 and HLE 36.8 vs Opus-5.0's 56.3
- Verbose and output-heavy at reasoning_effort=100 - 250M generated tokens on Artificial Analysis' index
- Vision reads documents, not diagnostics; no radiology or pathology claim
HIPAA-compliant hosting
DeepSeek-V4.1-Flash is available under our HIPAA-compliant design-partner program, with a signed Business Associate Agreement (BAA), encryption in transit and at rest, and access controls. We are currently onboarding design partners, with general availability coming soon. Run the model behind a unified OpenRouter-style API and swap to another model with a single parameter.
Pricing in context
Open source models like DeepSeek-V4.1-Flash can be served efficiently on optimized inference infrastructure, with savings passed through to you. Exact savings depend on the model and your volume, but open source inference is typically a fraction of the per-token cost of closed-source frontier models on Azure OpenAI or AWS Bedrock - without cloud egress lock-in or minimum commitments.
Design partner program
Run DeepSeek-V4.1-Flash under HIPAA
Join the waitlist to be prioritized. We'll reach out with a qualification call and early access.