Model launches
Introducing HIPAA-compliant GLM-5.3-Flash
GLM-5.3-Flash from Z.ai is a 320B MoE at 18B active, MIT-licensed, with hybrid sparse-plus-linear attention. Vision-capable, at flash-tier pricing.
TL;DR
GLM-5.3-Flash from Z.ai is the first natively multimodal GLM-5 model - 320B total, 18B active, MIT license, 1M context - now available under a signed BAA with zero data retention on our HIPAA-compliant design-partner program.
Healthcare teams evaluating GLM over the past two years have had a broader tradeoff: the flagship models are strong but expensive, the smaller ones are cheap but text-only. What has been missing is a vision-capable GLM at a genuinely deployable price point.
Today, that gap closes. GLM-5.3-Flash, Z.ai's first natively multimodal model in the GLM-5 series, is now available on our HIPAA-compliant API platform under a signed BAA with end-to-end encryption and strict zero data retention - with no proprietary vendor lock-in. It keeps the Q3-style speed-and-cost profile our readers have come to expect from a "flash" tier, while adding the vision capability that every prior GLM model in our catalog lacked.
What is GLM-5.3-Flash?
Released on August 25, 2026 to Hugging Face under the model ID
zai-org/GLM-5.3-Flash under the MIT license, GLM-5.3-Flash is a
320B total / 18B active Mixture-of-Experts model with a 1M token
context window on the hosted endpoint. Z.ai's model card describes a
hybrid attention architecture with sparse plus linear attention - a
first for the GLM series - plus Manifold-Constrained Hyper-Connections
(mHC) and a 30-trillion token multimodal pre-training corpus.
Native vision understanding comes built in from the start, not as a
bolt-on.
reasoning_effort is controllable with low, high, and max
levels, defaulting to max. The MIT license means there is no revenue
threshold and no attribution clause to read around, unlike the modified
MIT variants several frontier open models have been released under this
year.
Performance profile: what it's good at vs. what it's not
What GLM-5.3-Flash excels at
- Multimodal, not just vision. Z.ai's card describes GLM-5.3-Flash as a native multimodal release with a 30T-token pre-training corpus spanning text, images, and document layouts - a capability the previous GLM-5.2 flagship did not have at any price.
- Frontier-class reasoning breadth. On Z.ai's reported evaluation entries, it scores 63.4 on DeepSWE 1.1 (via mini-swe-agent, with 400K context) and 84.3 on Terminal-Bench 2.1 (Claude Code harness, 6-hour timeout) - agentic coding and terminal-work numbers in the Claude Opus and Kimi K3 band rather than the "budget slash" tier these labels usually imply.
- Real long-context, not just a listed spec. The 1M token context window is usable end-to-end on our serving layer, including on NL2Repo and other repo-level tasks that a smaller context window can make impractical.
- A clean license. MIT, the same class of licensing posture as DeepSeek-V4-Pro-0813, with no commercial-use restrictions to read before hosting.
- The
reasoning_effortworkhorse pattern. Defaulting tomaxfor reproduction-grade benchmarks while makinglowergonomic for high-volume chat means we can serve it flexibly for different healthcare workloads rather than treating "flash tier" as one fixed speed-versus-quality setting.
Limitations to keep in mind
- Vendor-reported benchmarks on a fast-moving harness. Z.ai's numbers for DeepSWE, Terminal-Bench, HLE, NL2Repo, and Toolathlon all use their own stated evaluation setups (different judges, different max generation lengths, different timeout windows). Treat single rows as directional, not as settled values, until third-party reruns land.
- "Flash" labels undersell it, which cuts both ways. Our own July post framing GLM-5.2 as a flagship was the correct read at the time; this release makes "Flash versus flagship" a less binary axis than it was two weeks ago. Teams picking purely by tier label will miss that this 18B-active model outperforms a 753B-parameter "flagship" Z.ai released a month ago.
- Not a DICOM or radiology tool. Without clinical validation, no "vision-capable open source model" in this catalog is a diagnostic instrument. Treat it as an expert reading assistant for charts, screenshots, scanned documents, and photos - nothing more.
- Pricing claims need context. Z.ai's "one-tenth the price" is a claim about tier positioning against GLM-5.2's hosting cost, not a per-token price we have validated end to end on our own infrastructure. The MIT license means we can reach that price, not that we have published a rate card for it yet. On our serving layer, GLM-5.3-Flash runs at $0.15 per 1M input, $0.03 per 1M cached input, and $0.50 per 1M output.
The benchmarks: how GLM-5.3-Flash compares
Benchmarks below are Z.ai-reported entries on the GLM-5.3-Flash model card (benchmarks evaluated by Z.ai via their stated harness configs - see the card's footnotes for exact methodology per benchmark; the '-' entries mark benchmarks not reported on that checkpoint by Z.ai).
| Benchmark (Focus) | GLM-5.3-Flash | Kimi K3 | DS-V4-Pro-0813 | GLM-5.2 |
|---|---|---|---|---|
| DeepSWE 1.1 (agentic coding) | 63.4 | 67.5 | 62.7 | 46.2 |
| Terminal-Bench 2.1 (terminal coding) | 84.3 | 88.3 | 87.9 | 81.0 |
| HLE (multidisciplinary reasoning) | 55.3 | - | 42.7 | - |
| Agents' Last Exam | - | 27.6 | 25.7 | 23.8 |
| AutomationBench (automation) | - | 30.8 | 31.8 | 12.9 |
| Toolathlon-Verified (tool use) | - | 76.5 | 74.1 | 59.9 |
| NL2Repo (repo-level generation) | - | - | 61.5 | 48.9 |
The direction of these rows, even with the caveats above, is clear: GLM-5.3-Flash sits in the same agentic-coding tier as DeepSeek-V4-Pro and Kimi K3 while staying roughly at Flash-class deployment cost, and the numbers Z.ai publishes hold up well against our own spot-checks of related checkpoints.
What the table does not capture is the context-window economics. Kimi K3 and DeepSeek-V4-Pro-0813 both headline 1M-token contexts; GLM-5.3-Flash matches that on the hosted endpoint, but we have not verified it against a third-party runner.
Why run GLM-5.3-Flash via our HIPAA-compliant API?
- Enterprise-grade privacy. Available under our HIPAA-compliant design-partner program with a signed BAA, end-to-end data encryption, and strict zero-data-retention on prompts and outputs - including images. Scanned documents, screenshots, and clinical paperwork are the data you cannot paste into consumer AI tools. Your patient data is never used for training. General availability is coming soon.
- Uncompromised performance. We handle the serving stack, the hybrid attention implementation, and the 1M-token context-window scaling - you get 320B-parameter multimodal reasoning without operating the hardware yourself.
- Transparent pricing. Under MIT with 18B active parameters, the per-token economics land far below closed multimodal frontier APIs on Azure OpenAI or AWS Bedrock, in the same Flash-class band as the DeepSeek family we already host.
Ready to put multimodal, Opus-class-adjacent reasoning under your BAA? Join the waitlist now.
Run it HIPAA-compliant
GLM-5.3-Flash on OpenMed Router
$0.15/M input · $0.50/M output · 1048576 context
Chris Williams, MD
Chris Williams, MD is a physician, clinical AI researcher and the co-founder of OpenMed Router, working to make open source AI models safely accessible to healthcare organizations under HIPAA. He writes about clinical AI, model selection, compliance, and the practical adoption of open source inference in clinical and operational workflows.
Join the waitlist