Best Open-Weight LLM for Enterprise in 2026: GLM-5.3 vs Kimi K3 vs DeepSeek V4 vs Qwen3.8 vs Hy4
The best open-weight LLM for enterprise deployment in 2026 is GLM-5.3-Flash for most organizations, GLM-5.3 for agentic coding teams with an 8-GPU node, and Kimi K3 for those with the cluster to feed it. That is the short answer after a summer that rewrote the category twice. In June and July, GLM-5.2 (Zhipu AI, MIT) and Kimi K3 (Moonshot AI, weights July 27) put last-generation-flagship intelligence on hardware you can own. Then, between August 10 and August 31, ten more open-weight models landed, five of them frontier-class MoE releases in the final week alone: GLM-5.3-Flash and GLM-5.3 from Z.ai, Qwen3.8-Flash-Next from Alibaba, Hy4-preview from Tencent, and DeepSeek-V4-Flash-Vision-Exp. Around them, DeepSeek-V4-Pro-0813 owns permissively licensed frontier reasoning, Qwen3.8-27B owns multimodal document work, and Meta's Muse Glimmer 30B and IBM's Granite 4.2 cover Western-origin procurement under Apache 2.0. Open weights are no longer one tier behind; on the published cards they sit inside three points of each other at the top, and the deciding factor has moved from capability to license and memory footprint.
What changed in September 2026?
This post was first published July 20, 2026 and refreshed September 3, 2026 after the August release wave. The material changes to our recommendations:
- GLM-5.2 is superseded. GLM-5.3 (August 28) shares the same 753B base with all gains from scaled post-training, reporting Terminal-Bench 2.1 of 88.2 versus a Terminal-Bench 3.0 jump from 4.6 to 28.3 over GLM-5.2, and it upgrades in place on existing vLLM and SGLang deployments. GLM-5.3-Flash (August 25-26) is a new 320B / 18B-active hybrid-attention base that beats GLM-5.2 on DeepSWE (63.4 vs 46.2) at less than half the memory. For most single-rack deployments, Flash is the new workhorse; GLM-5.3 is the coding upgrade.
- The GLM license split. GLM-5.3-Flash is plain MIT. GLM-5.3 ships under a custom "GLM-5.3 License" that adds a Z.ai security-review requirement for Model-as-a-Service operators with more than US$10 billion in group revenue. Embedding the model in an end-user product does not trigger it.
- Kimi K3 weights are shipped.
moonshotai/Kimi-K3has been live since about July 27 (2.8T total, 1.56 TB, MXFP4 supported, custom Kimi K3 License). It is no longer a pending release. - Two new permissively licensed frontier options. Tencent's Hy4-preview (770B / 49B, Apache 2.0, 1M context) and DeepSeek-V4-Pro-0813 (1.6-1.7T / 49B, MIT, 1M context) both land inside three points of K3 on Terminal-Bench 2.1.
- The single-GPU tier is real. Meta's Muse Glimmer 30B (Apache 2.0, August 10), IBM Granite 4.2 in 3B / 8B / 30B (Apache 2.0, August 25), and Qwen3.8-27B (Apache 2.0) give regulated teams frontier-derived models on a 24-32 GB card.
- Qwen3-VL is replaced in our stack by Qwen3.8-27B (dense, multimodal, Apache 2.0) and Qwen3.8-Flash-Next (125B / 6B active, multimodal, Qwen Community License 1.0).
The release-by-release detail, with a master table and hardware tiers, is in our August 2026 open-weight release roundup.
How do the 2026 open-weight flagships compare?
Figures are from the Hugging Face model cards and vendor documentation as of September 3, 2026.
| Model | Developer | Scale | Context | Modalities | License | Serving class | Best at |
|---|---|---|---|---|---|---|---|
| Kimi K3 | Moonshot AI | 2.8T MoE (104B active on card; press reports ~50B) | 1M | Text + native vision | Custom Kimi K3 License | Cluster (1.56 TB repo) | Long-horizon agentic runs; top card numbers |
| GLM-5.3 | Z.ai | 753B MoE, ~40B active | 1M | Text | Custom GLM-5.3 License (MIT-style + US$10B MaaS review clause) | 8-GPU node (~245-810 GB quantized) | Agentic coding; in-place GLM-5.2 upgrade |
| GLM-5.3-Flash | Z.ai | 320B MoE, 18B active | 1M (128K output) | Text, image, video, documents | MIT | 2-4 GPUs (109-200 GB GGUF) | Firm-wide workhorse; single-rack sovereignty |
| DeepSeek-V4-Pro-0813 | DeepSeek | 1.6-1.7T MoE, 49B active | 1M | Text | MIT | 8-GPU node (~893 GB repo) | Permissively licensed frontier reasoning; batch |
| DeepSeek-V4-Flash-Vision-Exp | DeepSeek | 304.6B MoE, 13B active | 1M | Text + vision | MIT | 4 GPUs (~168 GB repo) | Vision under MIT; document intake |
| Hy4-preview | Tencent | 770B MoE, 49B active | 1M | Text | Apache 2.0 | 8-GPU node (1.56 TB, TP=8) | Apache-licensed frontier alternative |
| Qwen3.8-Flash-Next | Alibaba | 125B MoE, 6B active (+51B N-gram table) | 262K native, 1M YaRN | Text, image, video | Qwen Community 1.0 | 2-4 GPUs (111 GB Q4, 173 GiB FP8) | Lowest decode cost; Qwen4 architecture preview |
| Qwen3.8-27B | Alibaba | 27B dense | 262K native, 1M extended | Multimodal | Apache 2.0 | Single GPU at 4-bit | OCR, forms, document understanding |
| Muse Glimmer 30B | Meta | ~29.6B dense (incl. vision encoder) | 131K+ | Text + vision | Apache 2.0 | Single 24-32 GB GPU (17 / 32 GB quants) | Western-origin edge deployment |
| Granite 4.2 (3B / 8B / 30B) | IBM | Dense | 131K shipped (512K claimed) | Text | Apache 2.0 | Single GPU | Western-origin procurement; native thinking modes |
| Mistral Large 3 | Mistral AI | 675B MoE, 41B active | Long | Text | Apache 2.0 | 8-GPU node | Western-origin frontier-class option (Dec 2025) |
No new Llama model shipped in the window; Muse Glimmer is Meta's current open-weight release.
What did the summer of 2026 change?
Three events in five weeks reset the enterprise calculus. On June 11, the US government ordered Anthropic to cut foreign access to Fable 5 and Mythos 5 within 48 hours — organizations outside the US that had standardized on rented frontier intelligence lost it overnight. On June 13, Zhipu released GLM-5.2 under MIT with no regional restrictions. On July 16, Moonshot announced Kimi K3, which independent first-day testing ranked third among all models — behind only the two proprietary flagships that had just become unavailable to much of the world — and shipped its weights on July 27.
August then finished the argument. Five labs — Moonshot, Z.ai, DeepSeek, Tencent, and Alibaba — now report Terminal-Bench 2.1 scores between 85.4 and 88.3 on open weights. The strategic conclusion writes itself: the newest proprietary models still lead some evaluations, but they are rented and revocable; the open bench is at parity on agentic coding and ownable. Weights on your disks cannot be un-shipped by an export order.
Which model wins each enterprise workload?
The firm-wide workhorse: GLM-5.3-Flash. MIT license, 320B total with 18B active so it decodes fast, natively multimodal, 1M-token context that swallows entire document productions, and a 4-bit GGUF of 200 GB (109 GB at 2-bit) that makes a two-GPU server genuinely sufficient. Z.ai states it approaches Claude Opus 4.8 on coding and agentic benchmarks at one-tenth the price. It is the first hybrid linear/sparse-attention GLM, which is why long-context KV cost drops sharply. See the GLM-5.3-Flash model page.
Agentic coding at the top of the open bench: GLM-5.3. Terminal-Bench 2.1 of 88.2, DeepSWE v1.1 of 66.9, text-only, same serving shape as GLM-5.2. The trade is memory (an 8-GPU node even at 2-4-bit) and a custom license that your counsel will read once and file. Our GLM-5.3 model page has the deployment detail; the older GLM-5.2 GPU requirements guide still describes the hardware, since the weights shape is unchanged.
Long-horizon autonomy at cluster scale: Kimi K3, if you can serve it. First-day independent testing verified 3+ hour autonomous runs completing 122 tasks from a single prompt, sub-agent swarms of 25 parallel verification agents, and vision-in-the-loop UI iteration. It is also, per that testing, the best open model at admitting uncertainty rather than hallucinating — decisive in legal, health, and financial work. The card reports Terminal-Bench 2.1 of 88.3 and GPQA Diamond of 93.5. The cost is hardware: a 1.56 TB repo and 2.8T parameters (our K3 hardware guide has the tiers).
Permissively licensed frontier scale: Hy4-preview and DeepSeek-V4-Pro-0813. Both fit an 8-GPU node, both report Terminal-Bench 2.1 above 85, and neither carries a bespoke clause: Hy4 is Apache 2.0, V4-Pro is MIT. DeepSeek remains the price-performance leader for overnight reasoning at scale, typically deployed beside the interactive workhorse to keep the batch queue off the primary rack.
Documents, OCR, and intake: Qwen3.8-27B, with DeepSeek-V4-Flash-Vision-Exp as the MIT option. Scanned records, engineering drawings, forms, and evidence photos need a vision model. Qwen3.8-27B (Apache 2.0, dense, multimodal, OSWorld-Verified 84.3) is the single-GPU standard; DeepSeek's V4-Flash-Vision-Exp (MIT, 13B active, up to 384 tokens per image) serves heavier pipelines from a 4-GPU node. Qwen3.8-Flash-Next adds a multimodal option whose 51B N-gram table can live in system RAM.
Procurement-constrained environments: Muse Glimmer, Granite 4.2, and Mistral Large 3. Where policy requires Western-origin weights, Meta's Apache 2.0 return (Glimmer, SWE-Bench Pro 51.2, 17 GB and 32 GB quants), IBM's Granite 4.2-30B (Apache 2.0, SWE-bench Verified 57.0), and Mistral Large 3 (Apache 2.0, 675B / 41B) are the answer. Worth stating plainly: self-hosted open weights of any origin are static files — checksummed, incapable of transmitting anything, air-gappable. Model origin is a governance choice, and sovereignty means you get to make it.
What actually decides the choice for a regulated enterprise?
Benchmarks rank models; deployments are chosen by constraints:
- Hardware budget. Single GPU → Qwen3.8-27B, Granite 4.2-30B, or Muse Glimmer. Two to four GPUs → GLM-5.3-Flash (+ Qwen3.8-Flash-Next). 8-GPU node → GLM-5.3, Hy4-preview, or DeepSeek-V4-Pro-0813. Cluster → add Kimi K3.
- License terms. With capability compressed into a three-point band, the license is now the differentiator. MIT (GLM-5.3-Flash, DeepSeek) and Apache 2.0 (Hy4, Glimmer, Granite, Qwen3.8-27B, Mistral Large 3) file cleanly; custom terms (GLM-5.3, Kimi K3, Qwen3.8-Flash-Next, Qwen3.8-Max) need a counsel note. We have not reviewed the full Qwen Community License text.
- Evidence obligations. OSFI E-23 (effective May 1, 2027), HIPAA, law-society duties — all reward a pinned, self-hosted model with a reproducible evaluation harness over any API whose behavior changes on a vendor's schedule. See how this plays out for financial services.
- Data gravity. If prompts contain client files, PHI, lending data, or controlled technical data, the model must come to the data — decades of CAD notes, tolerances, and failure analyses are exactly what an enterprise can never paste into a public chatbot, and exactly what a self-hosted model turns into an internal expert.
- The fine-tune. The largest capability gap in practice is not between open models — it is between a generic model and one fine-tuned on your own corpus. LoRA adapters trained in-environment on your precedents, tickets, and records routinely matter more than a few benchmark points, and they are portable across model upgrades.
The honest bottom line
There is no single best open-weight LLM — there is a best portfolio, and in September 2026 it looks like this: GLM-5.3-Flash as the sovereign workhorse, GLM-5.3 for the coding agents, Qwen3.8-27B for everything with pixels, DeepSeek-V4-Pro-0813 or Hy4-preview for the batch farm and permissively licensed frontier work, Kimi K3 where cluster-scale agentic capability justifies cluster economics, and Muse Glimmer, Granite 4.2, or Mistral Large 3 where origin policy binds. All of it runs on hardware you own, at a cost per token that falls as usage grows — and none of it can be revoked by a policy decision in someone else's capital. That last property is why the market is moving: Deloitte projects 70%+ of enterprises will run on-prem or edge AI by 2028. The models finally deserve the hardware.
Questions we get
Frequently asked questions
What is the best open-source LLM for business use in 2026?
GLM-5.3-Flash for most deployments as of September 2026: MIT license, 320B total / 18B active parameters, 1M-token context, native text-image-video input, Terminal-Bench 2.1 of 84.3, and a 200 GB 4-bit GGUF that serves from two H100/H200-class GPUs. Organizations running coding agents on an 8-GPU node step up to GLM-5.3 (Terminal-Bench 2.1 of 88.2, custom license); those with cluster-scale infrastructure and frontier workloads consider Kimi K3, Hy4-preview, or DeepSeek-V4-Pro-0813.
Is Kimi K3 the most powerful open-weight model?
On vendor-reported cards it still holds the top line: Terminal-Bench 2.1 of 88.3, DeepSWE of 67.5, GPQA Diamond of 93.5, with 2.8 trillion parameters, 1M-token context, and native vision. Its weights shipped July 27, 2026 under a custom Kimi K3 License in a 1.56 TB repo. The margin has collapsed, though: GLM-5.3 reports 88.2 on Terminal-Bench 2.1 and DeepSeek-V4-Pro-0813 87.9, and Hy4-preview reaches 85.4 under Apache 2.0. K3 is the leader for organizations that can serve it, not the default.
Can open-weight models replace GPT or Claude for enterprise work?
For most enterprise workloads in 2026, yes. Five labs now report Terminal-Bench 2.1 scores within three points of each other on open weights, and Z.ai states GLM-5.3-Flash approaches Claude Opus 4.8 on coding and agentic benchmarks at one-tenth the API price. Unlike the newest proprietary flagships, open weights cannot be revoked: the June 11, 2026 US export order that cut foreign access to Anthropic's Fable 5 and Mythos 5 does not touch weights already on your disks.
Which open-weight LLM should a regulated company choose?
The one whose evidence file you can build: a pinned, self-hosted model with reproducible evaluations and a license your counsel can file without exceptions. In practice that means GLM-5.3-Flash (MIT) as the workhorse, GLM-5.3 for coding agents after reviewing its custom license, Qwen3.8-27B or Qwen3.8-Flash-Next for document and vision pipelines, DeepSeek-V4-Pro-0813 (MIT) or Hy4-preview (Apache 2.0) for batch and frontier work on an 8-GPU node, and Muse Glimmer 30B or Granite 4.2 (both Apache 2.0) where procurement policy requires Western-origin weights.
Take the 40 Claude skills and the briefing with you
The Vault 2026 skills pack (calendar audits, hiring scorecards, calibration, continuity plans) plus the sovereignty briefing: model releases, deployment economics and regulatory shifts for regulated firms. One click to unsubscribe.
Ready to move from reading to running?
We design, build, fine-tune, host, and maintain sovereign AI deployments end to end.
Book a sovereignty assessment How deployment works