Model launches
Introducing HIPAA-compliant Kimi K3: frontier-class intelligence for healthcare
Kimi K3 - Moonshot AI's 2.8-trillion-parameter flagship with a 1M-token context window and native vision - is now available on our HIPAA-compliant design-partner program. #1 open-weight model on the Artificial Analysis Intelligence Index, under a BAA with zero data retention.
TL;DR
Kimi K3, the #1 open-weight model on the Artificial Analysis Intelligence Index (score 57, #3 overall behind only Claude Fable 5 and GPT-5.6 Sol), is now available HIPAA-compliant under a BAA - 1M-token context, native vision, and frontier reasoning at open-weights pricing.
When healthcare teams need frontier-level intelligence, they hit the same wall: the most capable models are locked behind proprietary APIs that are expensive, opaque, and hard to bring under HIPAA. The open-weight alternative used to mean compromising on capability. That trade-off just got a lot smaller.
Today, we are bridging that gap. Kimi K3, the new 2.8-trillion-parameter flagship from Moonshot AI, is now available on our secure, HIPAA-compliant API platform under our design-partner program - protected by a signed BAA and strict zero-data-retention, with no proprietary vendor lock-in.
K3 takes over the top spot in our Kimi lineup on the models page, with Kimi K2.7 Code moving into the general-purpose slot previously held by Kimi K2.6. You get the full Moonshot family at every price point: K3 for the hardest reasoning and agentic tasks, K2.7 Code for high-volume coding pipelines. See the full specs and capabilities →
What is Kimi K3?
Released on July 16, 2026 - with open weights following on July 27 - Kimi K3 is a 2.8-trillion-parameter sparse Mixture-of-Experts (MoE) model with 896 experts, of which only 16 are active per token. That sparse design delivers the cognitive depth of a model roughly 3.7x the size of GLM 5.2 while keeping per-token compute manageable.
What truly sets K3 apart is its architecture: Kimi Delta Attention, a hybrid
linear attention mechanism, paired with Attention Residuals. Together they
let the model efficiently hold a full 1-million-token context window - long
enough to reason over an entire repository, a full longitudinal patient record,
or weeks of agent activity - with native visual understanding (text and
image input) and always-on reasoning. There is no non-thinking mode: K3
always thinks, and you tune the depth of that thinking via reasoning_effort
(Low, High, or Max) rather than switching it off.
Performance profile: what it's good at vs. what it's not
Every model has its sweet spots and trade-offs. To help you orchestrate your medical or engineering pipelines, here is where Kimi K3 shines - and where it falls short.
What Kimi K3 excels at
- Long-horizon agentic workflows. K3 is built for autonomous execution - multi-step pipelines that run for hours or days without losing track of the goal. Community testers describe the jump from K2.6 to K3 as the most dramatic in the Kimi series.
- Repository-scale software engineering. With a 1M context window and always-on reasoning, it can digest entire codebases, execute autonomous refactors, and run deep debugging sessions - which is why health-tech engineering teams are already swapping it into their coding agents.
- Frontier-class reasoning. It ranks #2 overall on the Debate Benchmark, trailing only Claude Fable 5, and beats Claude Opus 4.8 and GPT-5.5 on coding and general-agent evaluations.
- Front-end and web development. K3 outperformed every competitor on Arena.ai's front-end web development benchmark.
- Native vision. It reads images as well as text - scanned records, forms, charts, and photos - opening up document-heavy clinical workflows without bolting on a separate vision pipeline.
Limitations to keep in mind
- Modest throughput. K3 generates roughly 34 tokens per second in independent testing - excellent for deep reasoning and batch work, but not for real-time, high-frequency chat. Pair it with a fast model like K2.7 Code for interactive tiers.
- Verbose generation. K3 "thinks out loud": it produced ~130M tokens across the Artificial Analysis Intelligence Index evaluation, so it consumes more output tokens per task than leaner models. Budget accordingly.
- Text and image only. Unlike K2.6, K3 has no native video input.
- A serious footprint to self-host. At 2.8T parameters, K3 needs roughly 1.7 TB of accelerator memory to run. As the r/LocalLLaMA community puts it: "does it really matter if it's open source if no one can run it?" — our hosted API removes that barrier entirely.
The benchmarks: how Kimi K3 compares
Kimi K3 doesn't just lead the open-weight pack; it is actively challenging the most expensive proprietary models on the market. On the independent Artificial Analysis Intelligence Index, K3 scores 57 - the highest of any open-weight model and #3 overall, behind only Anthropic's Claude Fable 5 and OpenAI's GPT-5.6 Sol, and ahead of Claude Opus 4.8 and GPT-5.5.
| Benchmark | Focus area | Kimi K3 performance |
|---|---|---|
| Artificial Analysis Intelligence Index | Aggregate multi-task intelligence | Score 57 - #1 open-weight, #3 overall (behind Claude Fable 5 & GPT-5.6 Sol; ahead of Claude Opus 4.8 & GPT-5.5) |
| Coding & general agents (Moonshot evals) | Agentic coding & tool use | Beat Claude Opus 4.8 and GPT-5.5 |
| Arena.ai front-end web development | Web & frontend | Outperformed all competitors |
| Debate Benchmark | Multi-turn reasoning | #2 overall, trailing only Claude Fable 5 |
Early community testing echoes the numbers: r/LocalLLaMA's consensus puts K3 at roughly 98% of Claude Fable 5's performance at a fraction of the cost - a gap that matters when you're running agents at scale.
Why run Kimi K3 via our HIPAA-compliant API?
K3 is open-weight, but deploying and serving a 2.8T MoE architecture internally - while maintaining rigid healthcare compliance standards - is an expensive, time-consuming engineering hurdle. By utilizing our API, you get the best of both worlds:
- Enterprise-grade privacy. Available under our HIPAA-compliant design-partner program with a signed Business Associate Agreement (BAA), end-to-end data encryption, and a strict zero-data-retention policy on your prompts and outputs. Your patient data is never used for training. General availability is coming soon.
- Uncompromised performance. We handle the hardware complexity, serving optimization, and context-window scaling - so you get frontier-class reasoning without needing a 1.7 TB accelerator cluster of your own.
- Transparent pricing. Kimi K3 runs at $3.00/M input, $0.30/M cached input, and $15.00/M output - a fraction of what Claude Fable 5 and GPT-5.6 Sol cost on Azure OpenAI and AWS Bedrock, with cache-hit pricing cutting input costs by ~90%.
Ready to put frontier-class reasoning under your BAA - for clinical reasoning, utilization review, or the next generation of health-tech coding agents? Join the waitlist now.
Run it HIPAA-compliant
Kimi K3 on OpenMed Router
$3/M input · $0.30/M cached input · $15/M output · 1048576 context
Chris Williams, MD
Chris Williams, MD is a physician, clinical AI researcher and the co-founder of OpenMed Router, working to make open source AI models safely accessible to healthcare organizations under HIPAA. He writes about clinical AI, model selection, compliance, and the practical adoption of open source inference in clinical and operational workflows.
Join the waitlist