Model bench · Zhipu AI / Z.ai

Deploy GLM-5.2 on-premise.

GLM-5.2 was the practical on-premise workhorse of mid-2026: a 744-billion-parameter mixture-of-experts with only 40B active parameters per token, a 1M-token context window, and a fully permissive MIT license. Released June 13, 2026 with no regional restrictions, it landed within a few points of Claude Opus 4.8 on key agentic benchmarks. As of September 2026 it is superseded: the same base model ships as GLM-5.3 with far stronger coding results, and the new GLM-5.3-Flash beats it across the published benchmarks in under half the memory while keeping the MIT license. Existing GLM-5.2 racks upgrade to GLM-5.3 by swapping checkpoints; most re-provision to Flash. This page remains as the sizing reference for deployments already running it.

Specifications

GLM-5.2 at a glance.

SpecGLM-5.2
Architecture744B-parameter mixture-of-experts (MoE), 40B active parameters per token
Context window1,000,000 tokens
ModalitiesText-only — pair with Qwen3-VL for vision workloads
LicenseMIT — fully permissive, no regional restrictions
ReleasedJune 13, 2026 · open weights immediately
Benchmark positionWithin a few points of Claude Opus 4.8 on key agentic benchmarks, at roughly one-fifth the cost
Serving classSingle-rack on-premise deployment — the 40B-active design is the enabler

The MIT license matters as much as the benchmarks: no usage restrictions, no acceptable-use gatekeeping, no license terms to renegotiate. The model is yours to run, fine-tune, and keep.

Hardware requirements

Sizing the deployment honestly.

Capabilities

What it’s best at.

GLM-5.2’s core strengths are software engineering and tool-driven agents: repository-scale coding, structured tool use, and reliable multi-step task execution. It is the model we deploy most — the copilot behind law-firm drafting, credit-union member service, and manufacturing knowledge systems — because it delivers near-flagship agentic quality at a hardware footprint a mid-sized organization can actually own.

The 1M-token context carries entire document productions, codebases, or policy manuals in a single pass. For vision workloads — scanned records, drawings, forms — we pair it with Qwen3-VL behind the same gateway.

Deployment patterns

How we deploy it.

The single-rack sovereign node. The signature deployment: one hardened rack on your premises or in in-country colocation, serving fine-tuned GLM-5.2 with SSO, role-based access, and full audit logging. Six to twelve weeks from assessment to production.

Fine-tuned house model. LoRA adapters trained on your precedents, templates, or archive — in your environment — make GLM-5.2 the model that speaks your firm’s language. MIT licensing keeps every derivative unambiguously yours. Air-gapped operation is fully supported.

Questions we get

Frequently asked questions

What GPU hardware does GLM-5.2 require?

The 40B-active MoE design is the key: although the full parameter count is 744B, per-token compute is that of a 40B model, which puts serving within reach of a single GPU rack. Quantized workgroup deployments start in the low-to-mid six figures of hardware; full-precision, firm-wide serving with long-context headroom fills a rack. Exact GPU counts depend on precision, context usage, and concurrency — the sovereignty assessment produces the sized bill of materials.

Is the MIT license really unrestricted for commercial use?

Yes. MIT is the most permissive mainstream license in software: commercial use, modification, fine-tuning, and internal deployment are all unambiguously permitted, with no regional restrictions and no revocation mechanism. Your fine-tuned adapters and every derivative remain your property.

How does GLM-5.2 compare to Kimi K3?

K3 is deeper — stronger on long-horizon agentic autonomy and equipped with native vision — but full-precision K3 is supercomputer-class, needing 64+ accelerators. GLM-5.2 lands within a few points of last-generation proprietary flagships on agentic benchmarks while fitting on one rack. Most organizations run GLM-5.2 as the daily workhorse and add K3 at the cluster or managed-facility tier only where workloads demand it.

The workhorse fits in one rack. Size yours.

The two-week sovereignty assessment sizes the hardware against your real workloads and hands you a written architecture with a cost model — before you buy a single GPU.

Book a sovereignty assessment