Model launches
Introducing HIPAA-compliant GLM-5.3
GLM-5.3, post-trained from the GLM-5.2 base, leads open models on Terminal-Bench 3.0 and CyberGym - now available under a signed BAA with zero data retention.
TL;DR
GLM-5.3, Z.ai's post-trained update on the same 753B MoE base as GLM-5.2, is now available under a signed BAA with zero data retention - the first open model to lead any frontier-tier terminal benchmark, and the first released under Z.ai's new security-review license.
Six weeks ago, our last GLM post described GLM 5.2 as the frontier class MoE flagship in our catalog - a 753B mixture-of-experts model that ran at roughly $1.4 per 1M input tokens, and was the top-scoring open model in the Artificial Analysis Intelligence Index at that moment. The frames have moved since. DeepSeek and Qwen shipped bigger releases later in August, and Z.ai followed with a checkpoint of their own that deserves the same careful read.
Today, we are adding GLM-5.3 - released August 14, 2026, weights published August 28, 2026 under a new license - to our catalog of HIPAA-compliant, BAA-backed inference under our design-partner program. It is text-only, it reuses the exact same base weights as GLM-5.2, and our honest read in this post is that it gets its gains in a very specific place: agentic coding under long-horizon, multi-turn evaluation. See it in context →
What is GLM-5.3?
GLM-5.3 is a 753B-parameter Mixture-of-Experts model, with the same MoE base as GLM-5.2 and every gain coming from post-training - Z.ai's card describes roughly one additional month of reinforcement learning on more executable environments, longer tasks, and stronger verifiers. It is a 753B total / native FP8 model with a 1M token context window on Z.ai's own endpoint, and unlike GLM-5.2 it is text-only - the vision capability our readers associate with the GLM 5 family rides on the separate GLM-5.3-Flash variant, which we covered separately earlier this week.
Z.ai released GLM-5.3 on August 14, 2026 under the tagline "Built to
Code. Ready for Cyber Defense." Unlike GLM-5.2, which shipped MIT-licensed
weights inside a week, this one held the weights for two weeks while
Z.ai ran what they describe as their most extensive risk review to date,
citing cyber capability that emerged faster than expected. The weights
landed on Hugging Face on August 28, exactly fourteen days later, under
a new glm-5.3 license rather than MIT - a custom license with a
separate clause for Model-as-a-Service operators that is worth reading
before you build around it.
That license deserves a proper note rather than a hand-wave, because GLM-5.2's MIT licensing set a bar the flagship release has dropped from:
License terms in plain English
Reading the actual GLM-5.3 License file on Hugging Face, three things
matter for running it in production. Two are the same everywhere, one is
not.
- Commercial use is permitted. The license text explicitly allows running, deploying, fine-tuning, modifying, and selling copies or derivative works, with no revenue cap.
- Attribution, not gating. A copyright and permission notice must be included in all copies, and the use must comply with applicable laws and regulations - standard modified-MIT-style conditions we have read many variants of this year.
- A security-review clause, not a revenue cap. If you run a "Model as a Service" (Z.ai defines this as third-party access to inference via an API, with meaningful control over inputs/parameters) and your affiliate-group aggregate revenue exceeds $10 billion over a rolling 12 months, you must pass a Z.ai security review before continuing commercially. The scope and method is at Z.ai's discretion. If you are a small health-tech vendor running inference for your own product through a design-partner program, this clause does not apply to you.
That last item is the notable gap versus MIT. It is a deliberately narrow clause aimed at very large Model-as-a-Service operators, and is a meaningfully different posture from the custom "qwen3.8-max" license on Alibaba's 2.4T release (which sets that threshold at $50M), so if you are licensing-comparing across catalogs, the GLM-5.3 wording is closer to MIT than the Qwen variant is.
Performance profile: what it's good at vs. what it's not
What GLM-5.3 excels at
- Coding and agentic work, genuinely across a spread of benchmarks. Terminal-Bench 2.1 at 88.2 (Claude Code 2.1.207 harness, max effort, temperature 1.0, top_p 1.0, max output 128K tokens) puts it ahead of DeepSeek-V4-Pro-0813 at 87.9, at parity with Kimi K3 at 88.3, and above Opus 4.8 at 85.0 - a real step past GLM-5.2's 81.0 on the same harness.
- First open model over a frontier tier on Terminal-Bench 3.0. At 28.3 it ranks above Kimi K3's 17.4 and Opus 4.8's 21.1 on the same third-party benchmark, though it still trails closed frontier at 33.7 (Fable-5 with fallback) and 34.6 (GPT-5.6 Sol). Worth noting that this row is Z.ai's first instance of an open model leading any frontier-tier terminal-benchmark field, which is a genuine shift rather than a leaderboard variance.
- Cyber capability at frontier level. CyberGym at 84.5 edges both Opus-4.8 (78.1) and Fable-5 with fallback (83.8) on the same card table - and on AutomationBench, GLM-5.3 posts 48.2 against Opus-4.8's 41.0 and Kimi K3's 46.7, the largest gain in the release.
- Longer-horizon software engineering. DeepSWE v1.1 climbs from 46.2 (GLM-5.2) to 66.9; FrontierSWE reaches 78.1 against Opus-4.8's 66.5 on the same evaluation.
- GDPval-AA v2 Elo 1769, ahead of both Fable-5 (1743) and GPT-5.6 Sol (1730) on the vendor's own agentic real-world work benchmark, and Toolathlon Verified at 73.0.
Limitations to keep in mind
- It cannot see. This is a text-only release; vision work goes through GLM-5.3-Flash or another vision-capable model. If your workflow starts from an image, scan, or screenshot, this is not the entry point.
- Cyber benchmarks and exploitation gaps. Z.ai frames GLM-5.3 as defensive cybersecurity; on ExploitBench it posts 54.4 against Fable-5's 78.0 and GPT-5.6 Sol's 76.5. On ExploitGym (2h / 6h) it sits at 105 / 130 against Fable-5's 181 / 247. The aggressive-security side of the release is real but not leading. For readers thinking about security-adjacent workloads, that distinction matters and is exactly the one Z.ai is drawing with their "defense, not offensive exploit generation" framing.
- The upside came from post-training, not architecture. There is no new pre-training run here - same 753B MoE base as GLM-5.2, same MoE routing, same hybrid-attention shape. Every number on the model card is a post-training gain; the architecture story in this release is conservative, not novel.
- A non-standard license, but a narrow clause. This is the first GLM 5 checkpoint not shipped under MIT. Reading the license file directly: commercial use, deployment, and derivative works are explicitly permitted with no revenue cap, and the only non-standard clause - "Model as a Service" operators (third-party access to inference via an API) and their affiliates exceeding $10 billion in aggregate revenue per rolling 12 months must pass a Z.ai security review. The $10B threshold is deliberately far above where any single health-tech vendor or design-partner-stage inference provider operates, so the practical effect for teams in this space is minimal, and the clause is closer to MIT than the custom Qwen licenses this season.
- The headline leads with the right strengths. Terminal-Bench 3.0 leadership is real, but it is one benchmark row in a table where DeepSWE moved from 46.2 to 66.9 and CyberGym moved 77.2 to 84.5, both as post-training effects. Beat-benchmark headlines flatten that context.
- A gated release timeline. Unlike GLM-5.2's near-immediate MIT release, the model took a deliberate two-week gated release path - if your compliance or procurement cycle requires vendor documentation on licensing posture before you onboard a checkpoint, this is the first GLM release where the review shows up in the timeline rather than in the license file.
Model-card context worth reading
The GLM-5.3 model card explicitly notes that Z.ai's benchmark tables are BEST published scores across harnesses - their tables mix self-run harness numbers with the best published score from other sources for baseline models. Cross-checking our own posts, the official Z.ai Terminal-Bench 2.1 figure (88.2) is consistent with how the Artificial Analysis Terminus inference matches that figure, and the numbers we report here are as the vendor published them.
The benchmarks: how GLM-5.3 compares
Benchmarks are from the official Z.ai GLM-5.3 model card on Hugging Face, with GLM-5.3's own runs at Z.ai-documented harness settings (Claude Code 2.1.207, max effort, temperature 1.0, top_p 1.0, 128K max output, 1M context window).
| Benchmark (Focus) | GLM-5.3 | GLM-5.2 | Kimi K3 | DS-V4-Pro-0813 | Opus 4.8 |
|---|---|---|---|---|---|
| Terminal-Bench 2.1 (agentic coding) | 88.2 | 81.0 | 88.3 | 87.9 | 85.0 |
| Terminal-Bench 3.0 (harder terminal) | 28.3 | 4.6 | 17.4 | - | 21.1 |
| DeepSWE v1.1 (agentic software eng) | 66.9 | 46.2 | 67.5 | 62.7 | 58.0 |
| NL2Repo (code generation) | 58.0 | 48.9 | 58.0 | 61.1 | 69.7 |
| CyberGym (security) | 84.5 | 77.2 | 80.0 | 83.3 | 78.1 |
| ExploitBench (security, offensive) | 54.4 | 24.4 | 32.2 | - | 40.0 |
| Toolathlon Verified (tool use) | 73.0 | 59.9 | 76.5 | 74.1 | 76.2 |
| AutomationBench (agentic automation) | 48.2 | 26.2 | 46.7 | 43.2 | 41.0 |
| Agents' Last Exam (agentic tasks) | 28.5 | 23.8 | 27.6 | 25.7 | 25.7 |
| HLE with Tools (multi-add reasoning) | 62.5 | 54.7 | 59.8 | 60.0 | 57.9 |
| GDPval-AA v2 (agentic knowledge) | 1769 | 1508 | 1682 | 1590 | 1588 |
Reading the table honestly: on Terminal-Bench 3.0, DeepSWE, CyberGym, AutomationBench, Agents' Last Exam, and GDPval-AA v2, GLM-5.3 is the top open-weight entry in the row. On ExploitBench, still very much a frontier's-game benchmark, it lags well behind closed frontier. And the rows where it is crowded in at the top are the exact shapes of most healthcare coding work: long-horizon task execution, more tool calls per task, and higher bars for verifiable output than either GLM-5.2 or the previous open-weight entries could reliably deliver.
What this means for healthcare work
Three concrete uses, based on where the benchmark gains actually land:
- Clinical-backend automation. The AutomationBench climb (26.2→48.2) plus Agents' Last Exam improvement is exactly the profile you want for multi-step automation in health-adjacent back-office workflows: eligibility, prior-auth packet assembly, referral routing, claim-support documentation.
- Long-horizon engineering tasks across wide context. With a 1M context on the served endpoint, long-horizon engineering effort over 100k+ code files is realistic, which makes GLM-5.3 a credible option for clinical-data pipelines or internal tooling maintenance alongside your EHR's own engineering work.
- Cyber on the defensive side. GLM-5.3's CyberGym result is what Z.ai highlights, and for teams evaluating defensive security posture (vulnerability discovery, security-review automation), it holds up against closed frontier models in the same row.
Why run GLM-5.3 via our HIPAA-compliant API?
- Enterprise-grade privacy. Available under our HIPAA-compliant design-partner program with a signed BAA, end-to-end data encryption, and a strict zero-data-retention policy on your prompts and outputs. Your patient data is never used for training. General availability is coming soon.
- Uncompromised performance. We handle the 753B MoE serving stack and 1M-token context-window scaling - you get frontier-class agentic performance without operating a multi-GPU cluster.
- Transparent pricing. GLM-5.3 runs at $1.40 per 1M input tokens and $4.40 per 1M output tokens on the underlying serving layer - the same Z.ai published rate as GLM-5.2, for a materially stronger checkpoint underneath.
Ready to run frontier-class open-source coding under your BAA? Join the waitlist now.
Run it HIPAA-compliant
GLM-5.3 on OpenMed Router
$1.4/M input · $4.4/M output · 1048576 context
Chris Williams, MD
Chris Williams, MD is a physician, clinical AI researcher and the co-founder of OpenMed Router, working to make open source AI models safely accessible to healthcare organizations under HIPAA. He writes about clinical AI, model selection, compliance, and the practical adoption of open source inference in clinical and operational workflows.
Join the waitlist