← Back to blog

Model launches

Introducing HIPAA-compliant DeepSeek-V4-Pro-0813

DeepSeek-V4-Pro-0813, the GA checkpoint of DeepSeek's 1.6T flagship, scores 87.9 on Terminal-Bench and is available on our HIPAA-compliant design-partner program.

Chris Williams, MD

TL;DR

DeepSeek-V4-Pro-0813, the official GA checkpoint of DeepSeek's 1.6T parameter flagship, is now available under a signed BAA with zero data retention - a model that went from losing to its own Flash sibling on agentic benchmarks to an 87.9 on Terminal-Bench 2.1.

Healthcare teams building on large language models keep hitting the same wall. The closed frontier APIs that are good enough for clinical work are expensive at scale, and sending patient data to them inside a business associate agreement is only possible with vendors who offer one. Open weights solved the cost problem over the past year, but the best open agentic model and the best hosted one were rarely the same model.

Today, we are bridging that gap. DeepSeek-V4-Pro-0813, the official general availability release of DeepSeek's 1.6 trillion parameter flagship, is now available on our HIPAA-compliant design-partner program - under a signed BAA, with end-to-end encryption and strict zero data retention on your prompts and outputs. There is no proprietary vendor lock-in and no training on your data.

This post is the checkpoint update to our July review of DeepSeek-V4-Pro. If you evaluated the preview in the spring and moved on, the GA numbers below are worth a second look.

What is DeepSeek-V4-Pro-0813?

DeepSeek released DeepSeek-V4-Pro-0813 on August 13, 2026, less than two weeks after Flash-0731, as the general availability release of the V4 Pro preview. The architecture is unchanged: a 1.6T parameter Mixture-of-Experts (MoE) activating roughly 49B parameters per token, with a 1M token context window and the V4 family's hybrid attention (CSA + HCA, cutting KV cache to about 10% of a standard footprint), plus Manifold-Constrained Hyper-Connections for signal stability at depth.

Two things did change. First, DeepSeek attached its DSpark speculative decoding module to the served checkpoint, which improves throughput on agentic workloads without a separate draft model. Second, the reasoning_effort parameter now exposes three levels - low, high, and max - with DeepSeek recommending up to 384K output tokens at the high and max levels. The weights ship under the MIT license, the cleanest licensing posture of any frontier-class model we host.

Performance profile: what it's good at vs. what it's not

What DeepSeek-V4-Pro-0813 excels at

  • Closing the agentic gap. The preview scored 72.1 on Terminal-Bench 2.1 and just 12.8 on DeepSWE. The GA checkpoint scores 87.9 and 62.7 respectively on the same harnesses (vendor-reported, DeepSeek Harness minimal mode, max effort) - the preview's worst weakness is now one of its strengths.
  • Production agentic coding. At 61.5 on NL2Repo-Bench and 83.3 on CyberGym, it sits within a few points of Claude Opus-4.8 (69.7 and 78.3) at roughly a fifth of the price.
  • Long-horizon automation. AutomationBench (Public) improved from 12.8 to 31.8. For back-office work - eligibility checks, referral intake, prior-auth packet assembly - the failure modes that made the preview unreliable in long chains have largely been repaired.
  • Clinically adjacent reasoning at frontier level. It holds 74.1 on Toolathlon-Verified and 25.7 on Agents' Last Exam, both within a point or two of Kimi K3 and Opus-4.8.
  • A clean license. MIT means no revenue threshold, no attribution clause, and no negotiation before you offer it as part of a hosted service to your own customers. For a health-tech company this removes an entire legal review from your deployment path.

Limitations to keep in mind

  • Text-only. It cannot natively process images, charts, scanned documents, or video. If your workflow starts from a DICOM image, a screenshot, or a chart in a scanned chart-note, this is the wrong model on its own - pair it with a vision-capable model in your pipeline.
  • Slow at max effort. DeepSeek recommends long output budgets at high and max reasoning effort, and measured output speed is in the ~36-60 tokens per second band depending on the serving stack. For interactive triage UIs, route to a faster model and keep V4-Pro-0813 for background, long-horizon tasks where "correct eventually" beats "fast and wrong".
  • 1.6T parameters are not self-hostable for most teams. The full precision footprint requires multi-node GPU infrastructure. Unless you operate a data center, the practical deployment path is a managed API, which is exactly the constraint our design-partner program addresses.
  • A moving target upstream. DeepSeek has signalled a V4.1 generation, and served checkpoints may shift; if your evaluation is sensitive to exact checkpoint identity, pin the -0813 name in your integration.

The benchmarks: how DeepSeek-V4-Pro-0813 compares

Benchmarks below are from the official DeepSeek model card on Hugging Face (vendor-reported, evaluated with DeepSeek Harness minimal mode at max reasoning effort, temperature 1.0, top_p 0.95).

Benchmark (Focus)DS-V4-Pro-0813DS-V4-Flash-0731GLM 5.2Kimi K3Opus-4.8
Terminal-Bench 2.1 (terminal coding)87.982.781.088.385.0
DeepSWE (software engineering)62.754.446.267.558.0
Cybergym (security reasoning)83.376.7-80.078.3
NL2Repo (repo-level generation)61.554.248.9-69.7
Toolathlon-Verified (tool use)74.170.359.976.576.2
Agents' Last Exam (agentic mid)25.725.223.827.625.7
AutomationBench (automation)31.825.112.930.827.2

On the independent Artificial Analysis Intelligence Index v4.3, the model scores around 36 at max effort - roughly eleventh place overall and among the cheapest frontier-class models measured on cost per task.

Why run DeepSeek-V4-Pro-0813 via our HIPAA-compliant API?

  • Enterprise-grade privacy. Available under our HIPAA-compliant design-partner program with a signed BAA, end-to-end data encryption, and a strict zero-data-retention policy on your prompts and outputs. Your patient data is never used for training. General availability is coming soon.
  • Uncompromised performance. We handle the hardware complexity, serving optimization including DSpark speculative decoding, and context-window scaling - you get frontier-class agentic coding without operating multi-node accelerator clusters yourself.
  • Transparent pricing. DeepSeek-V4-Pro-0813 runs at $1.32 per 1M input tokens and $3.96 per 1M output tokens on the underlying serving layer - a fraction of Claude Opus-class pricing on Azure OpenAI or AWS Bedrock, with no egress lock-in and no minimum commitment.

Ready to run clinical-grade agentic coding under your BAA? Join the waitlist now.

Run it HIPAA-compliant

DeepSeek-V4-Pro on OpenMed Router

$1.32/M input · $3.96/M output · 1048576 context

Chris Williams, MD

Chris Williams, MD is a physician, clinical AI researcher and the co-founder of OpenMed Router, working to make open source AI models safely accessible to healthcare organizations under HIPAA. He writes about clinical AI, model selection, compliance, and the practical adoption of open source inference in clinical and operational workflows.

Join the waitlist

Be first in line for HIPAA-compliant open source inference