Qwen
Qwen3.8-Max
Specs
- Input price
- $2/M
- Cached input price
- $0.25/M
- Output price
- $6/M
- Context window
- 262.144K tokens
The first Qwen-Max-class open-weight release - 2.4T parameters with 95B active per token, frontier agentic coding at open-weights economics.
Capabilities
- 2.4T parameter MoE (512 experts, 10 routed + 1 shared per layer)
- 92-layer stack of Gated DeltaNet and Gated Attention, context extensible to ~1M
- Terminal-Bench 2.1 at 86.6 and PaperBench 93.0 - frontier research reasoning
- reasoning_effort (xhigh, medium, low) plus preserve_thinking control
Best for
- Long, multi-step research and evidence synthesis (PaperBench 93.0)
- Health and legal reasoning: HealthBench 60.2, PLawBench 73.2
- Frontier-tier terminal and agentic coding without closed-API pricing
Limitations to keep in mind
- Open checkpoint is text-only and always-on reasoning - verbosity costs at volume
- Full BF16 checkpoint targets GB300 NVL72-class racks; self-hosting is unrealistic for most teams
- Custom qwen3.8-max license: attribution plus MaaS review above $50M rolling revenue
HIPAA-compliant hosting
Qwen3.8-Max is available under our HIPAA-compliant design-partner program, with a signed Business Associate Agreement (BAA), encryption in transit and at rest, and access controls. We are currently onboarding design partners, with general availability coming soon. Run the model behind a unified OpenRouter-style API and swap to another model with a single parameter.
Pricing in context
Open source models like Qwen3.8-Max can be served efficiently on optimized inference infrastructure, with savings passed through to you. Exact savings depend on the model and your volume, but open source inference is typically a fraction of the per-token cost of closed-source frontier models on Azure OpenAI or AWS Bedrock - without cloud egress lock-in or minimum commitments.
Design partner program
Run Qwen3.8-Max under HIPAA
Join the waitlist to be prioritized. We'll reach out with a qualification call and early access.