← Back to models
DeepSeek logo

DeepSeek

DeepSeek-V4.1-Flash

Specs

Input price
$0.30/M
Cached input price
$0.006/M
Output price
$1.20/M
Context window
1M tokens

DeepSeek's Causal Encoder-Decoder debut - a 552B multimodal MoE with 8B/16B asymmetric activation, native vision, and a KV cache roughly a quarter of V4-Flash's.

Capabilities

  • 552B MoE (1 shared + 384 routed experts, 6 active per token) split into a 20-layer causal encoder and 20-layer decoder
  • Terminal-Bench 2.1 at 90.6 (DeepSeek Harness Minimal, reasoning_effort=100) - top of the model card's comparison table
  • AutomationBench 54.8 and Agent's Last Exam 31.8 lead the table; CyberGym 88.1 is the highest security score
  • 890 bytes per token global KV cache (FP4 main KV, SWA Bounded Replay) at a 1M-token context window
  • Native vision from pre-training: DocVQA 95.6, Chartography with tools 78.9, BabyVision with tools 89.6

Best for

  • Input-heavy clinical document pipelines: prior-auth packets, chart exports, and claims history
  • Long-horizon automation agents: claim assembly, referral routing, and back-office workflows
  • Agentic coding and security-adjacent review at Flash-class cost

Limitations to keep in mind

  • Trails the closed frontier on the hardest suites: Terminal-Bench 3.0 at 30.0 and 4.0 at 31.2
  • Deep reasoning is a tier behind: GPQA Diamond 90.9 and HLE 36.8 vs Opus-5.0's 56.3
  • Verbose and output-heavy at reasoning_effort=100 - 250M generated tokens on Artificial Analysis' index
  • Vision reads documents, not diagnostics; no radiology or pathology claim

HIPAA-compliant hosting

DeepSeek-V4.1-Flash is available under our HIPAA-compliant design-partner program, with a signed Business Associate Agreement (BAA), encryption in transit and at rest, and access controls. We are currently onboarding design partners, with general availability coming soon. Run the model behind a unified OpenRouter-style API and swap to another model with a single parameter.

Pricing in context

Open source models like DeepSeek-V4.1-Flash can be served efficiently on optimized inference infrastructure, with savings passed through to you. Exact savings depend on the model and your volume, but open source inference is typically a fraction of the per-token cost of closed-source frontier models on Azure OpenAI or AWS Bedrock - without cloud egress lock-in or minimum commitments.

Design partner program

Run DeepSeek-V4.1-Flash under HIPAA

Join the waitlist to be prioritized. We'll reach out with a qualification call and early access.