The model bench
Flagship open-weight models, deployed as of September 2026.
Five frontier-class open-weight releases landed in nine days at the end of August 2026 — GLM-5.3-Flash, Qwen3.8-Flash-Next, GLM-5.3, Tencent Hy4 and DeepSeek V4-Flash-Vision. These are the models we deploy today — selected, quantized, and served to match your workload and governance requirements, on infrastructure you own.
| Model | Scale | Context | Strengths | License / weights |
|---|---|---|---|---|
| GLM-5.3-Flash Zhipu AI / Z.ai | 320B MoE · 18B active — 4-bit fits in ~200 GB; the new single-node workhorse | 1M tokens · native vision | General assistant, agentic coding and tool use, document/image/video understanding — one checkpoint replaces the old GLM-5.2 + vision-model pair. Deployment guide → | MIT · Aug 25–26, 2026 |
| GLM-5.3 Zhipu AI / Z.ai | 753B MoE · ~40B active — in-place upgrade of GLM-5.2 | 1M tokens · text-only | The strongest open agentic-coding model: Terminal-Bench 2.1 88.2, DeepSWE 66.9, security research. For engineering organizations. Deployment guide → | GLM-5.3 License · Aug 28, 2026 |
| Kimi K3 Moonshot AI | 2.8T MoE · 104B active — the largest open model released | 1M tokens · native vision | Deep multi-hour autonomous runs, vision-in-the-loop work, document & spreadsheet analysis. Supernode-class hardware. Deployment guide → | Kimi K3 License · weights Jul 27, 2026 |
| DeepSeek V4 DeepSeek | V4-Pro-0813: 1.6–1.7T MoE · 49B active · V4-Flash: 284B · 13B active | 1M tokens · Flash-Vision variant adds images | Cost-efficient frontier reasoning at scale; the batch engine for high-volume analysis, extraction and classification. Deployment guide → | MIT |
| Qwen3.8 family Alibaba | Flash-Next: 125B · 6B active (+51B N-gram table in RAM) · 27B dense · 2.4T Max | 262K native → 1M · multimodal | Efficiency and document/vision work: Flash-Next runs in ~75–110 GB quantized; Qwen3.8-27B is the Apache-2.0 workstation model. Deployment guide → | Qwen Community 1.0 / Apache 2.0 (27B) |
| Llama · Mistral Meta / Mistral AI | Multiple scales | Long context | Western-origin open weights for organizations whose procurement or security policy requires them. Deployment guide → | Open weights |
On model origin: self-hosted open weights are static files — audited, checksummed, and incapable of transmitting anything. No data flows back to the lab that trained them, and fully air-gapped operation is available. Where policy requires Western-origin models regardless, we deploy them. Sovereignty means the choice is yours.
On honest sizing: full-precision Kimi K3 is supercomputer-class — Moonshot itself recommends 64+ accelerators. Most organizations don’t need that. GLM-5.3-Flash’s 18B-active design puts a frontier-class, multimodal, MIT-licensed model on a single node (about 200 GB at 4-bit); the full GLM-5.3 fills an 8-GPU node; K3 is available through large-cluster builds or our managed sovereign facility tier. The assessment tells you which tier your workloads actually justify.
On licenses: licensing is now the differentiator. GLM-5.3-Flash, DeepSeek V4 and Mistral Large 3 are MIT or Apache 2.0; Tencent Hy4, IBM Granite 4.2, Meta Muse Glimmer and Qwen3.8-27B are Apache 2.0; GLM-5.3, Kimi K3, Qwen3.8-Max and Qwen3.8-Flash-Next ship under custom licenses that are fine for internal enterprise use but need a procurement read. Every deployment file we hand over includes the license analysis.
Start with the sovereignty assessment.
Two weeks, fixed fee. We map your data obligations and workloads, size the hardware, and hand you a written architecture with a real cost model — whether or not you build with us.
Book a sovereignty assessment