Qwen3.8-Flash-Next Hardware Requirements for Self-Hosting
Qwen3.8-Flash-Next self-hosting: 125B MoE, 6B active, FP8 at 173 GiB, 4-bit GGUF at 111 GB, N-gram table in system RAM, and the Qwen Community License.
Field notes
Hardware requirements, deployment economics, and the compliance landscape — PIPEDA, Québec Law 25, PHIPA, OSFI Guideline E-23HIPAA, GLBA, CMMC, ITAR, and state privacy acts — written by the team that builds sovereign AI infrastructure for a living.
Qwen3.8-Flash-Next self-hosting: 125B MoE, 6B active, FP8 at 173 GiB, 4-bit GGUF at 111 GB, N-gram table in system RAM, and the Qwen Community License.
OpenClaw 2.0 (v2026.8.1) vs Hermes Agent v0.21.0: local-model support, SQLite migration, Bot Mode, sandboxing, and how to run either fully on-prem.
Ten open-weight LLMs shipped in August 2026: GLM-5.3, GLM-5.3-Flash, Qwen3.8-Flash-Next, Hy4, DeepSeek V4. Params, context, licenses, and hardware tiers.
GLM-5.3 on-prem guide: 753B MoE, ~40B active, 1M context, Terminal-Bench 2.1 88.2. Memory by precision, node layouts, and the new license clause explained.
GLM-5.3-Flash sizing guide: 320B MoE, 18B active, multimodal, 1M context, MIT license. Quant-by-quant memory table (93–642 GB), node layouts, what it replaces.
Chels.ai launches: we design, build, fine-tune, host, and maintain on-premise open-weight AI — Kimi K3, GLM-5.2 — for regulated firms in Canada and the US.
September 2026 comparison: GLM-5.3-Flash, GLM-5.3, Kimi K3, DeepSeek V4, Hy4, Qwen3.8, Granite 4.2, Muse Glimmer. Hardware, licensing, which to deploy.
Kimi K3 self-hosting guide: 2.8T MoE, 1M-token context, ~1.4TB native-4-bit checkpoint, deployment tiers, and Moonshot's 64+ accelerator guidance.
OSFI Guideline E-23 puts AI models under full model-risk management by May 1, 2027. What it requires — and why self-hosted models make the file defensible.
Deploy an LLM with zero network connectivity: CMMC and ITAR requirements, the architecture, and the telemetry and license-check traps that break isolation.
How Quebec Law 25 applies to AI tools: cross-border assessments, automated-decision transparency, and penalties up to C$25M or 4% of worldwide turnover.
GLM-5.2 deployment guide: 744B MoE with 40B active parameters, MIT license, 1M context — precision tiers, GPU counts, and single-rack serving that works.
Five architectures for using LLMs with PHI under HIPAA — and why a fully on-premise deployment needs no business associate agreement at all.
The US CLOUD Act (18 USC 2713) reaches data held by US providers even in Canadian regions. What that means for Canadian AI deployments — and the fix.
What on-premise LLM deployment actually costs in 2026: hardware tiers from a single node to a Kimi K3 cluster, and the ~2M tokens/day break-even vs APIs.
US v. Heppner (SDNY, Feb 2026) held consumer-AI chats are not privileged. What lawyers in Canada and the US can defensibly do with client files and AI.
Two weeks, fixed fee. We map your data obligations and workloads, size the hardware, and hand you a written architecture with a real cost model — whether or not you build with us.
Book a sovereignty assessment