Model bench · Alibaba
Deploy Qwen3.8 on-premise.
The Qwen3.8 family, released through August 2026, is the efficiency corner of the open-weight bench. Qwen3.8-Flash-Next (August 26) is a 125-billion-parameter mixture-of-experts with just 6B active parameters, a 51B N-gram embedding table designed to live in system RAM rather than VRAM, a 4B multi-token-prediction head, native text-image-video input, and 262K context extensible to 1M — an early preview of the Qwen4 architecture. Qwen3.8-27B is the dense, Apache-2.0, multimodal workstation model. Both read what your organization actually holds — scanned records, drawings, forms, evidence — without a page leaving your network.
Specifications
Qwen3.8 at a glance.
| Spec | Qwen3.8 |
|---|---|
| Qwen3.8-Flash-Next | 125B MoE · 6B active · 512 experts (10 routed + 1 shared) · 51B N-gram table + 4B MTP head · hybrid Gated DeltaNet + Qwen Sparse Attention · 48 layers |
| Flash-Next modalities & context | Text, image, video in; text out · 262,144 native → 1,000,000 with YaRN |
| Flash-Next sizes | FP8 172.78 GiB · BF16 335.28 GiB · Unsloth GGUF: UD-Q4_K_XL 111 GB · UD-Q2_K_XL 78.9 GB · UD-IQ1_S 72.5 GB · Qwen/Qwen3.8-Flash-Next |
| Flash-Next benchmarks (card) | SWE-bench Pro 62.5 · DeepSWE 58.7 · GPQA Diamond 91.7 · LiveCodeBench v6 91.9 · CoWorkBench 73.9 · MathVision 95.7 |
| Flash-Next license | Qwen Community License 1.0 — not Apache 2.0; procurement should read the terms before commercial redistribution |
| Qwen3.8-27B | 27B dense · multimodal · 262K → 1M · Apache 2.0 · SWE-bench Pro 61.7 · Terminal-Bench 2.1 73.0 · Qwen/Qwen3.8-27B |
| Qwen3.8-Max (2.4T / 95B active) | Open weights under a custom license (Qwen/Qwen3.8-2.4T-A95B) · text-only · 1M context — cluster-tier only |
| Serving | vLLM · SGLang · TokenSpeed · KTransformers · llama.cpp · Transformers; TP4 in the reference configs, TP2 minimum validated FP8 on GB300 |
The N-gram table is the genuinely new idea: 51B parameters of bigram/trigram embeddings that “can be offloaded to host memory and overlapped with model computation.” That moves a large share of the model into cheap system RAM — a new axis for single-node sizing that no other frontier-class model currently offers.
Hardware requirements
Sizing the deployment honestly.
Workstation & workgroup tier
Qwen3.8-27B at 4-bit fits a 24–32 GB accelerator, and Flash-Next at 2–4 bits runs in roughly 75–110 GB of accelerator memory with the N-gram table in system RAM. This is the tier for departmental document pipelines, intake and OCR-style extraction, and evaluation before a larger commitment.
FP8 node tier
Flash-Next FP8 (about 173 GiB) on a four-GPU node — TP4 in the reference configuration, TP2 the minimum validated FP8 layout on GB300-class parts — serves multimodal, 1M-context workloads firm-wide. With only 6B active parameters, decode is fast and tokens per dollar are the best on the bench.
Capabilities
What it’s best at.
Qwen3.8 earns its place on regulated benches because the most sensitive material is so often paper and pixels: patient charts, discovery productions, engineering-change orders, handwritten notes, screen recordings. Flash-Next reads documents, images and video natively, extracts structure, and reasons over it in one pass — then hands clean output to whatever language model sits beside it. Its CoWorkBench (73.9) and AndroidWorld (84.5) scores make it a strong agent runtime for tool-driven back-office work.
Against GLM-5.3-Flash — the other efficiency flagship of late August — Flash-Next is smaller and cheaper to run but carries the Qwen Community License rather than MIT and a 262K native context. Where the license must be permissive with no reading required, Qwen3.8-27B (Apache 2.0) or GLM-5.3-Flash (MIT) are the choices; where every gigabyte of VRAM matters, Flash-Next is. We benchmark on your documents before choosing.
Deployment patterns
How we deploy it.
The intake pipeline. Documents enter, structured data leaves — inside your network. Flash-Next performs understanding and extraction, a language model classifies and routes, and nothing transits a third party at any step.
Multimodal retrieval. Paired with GLM-5.3 or GLM-5.3-Flash over a PLM, DMS, or records archive, Qwen3.8 indexes what scanners captured and text search never could — making decades of paper conversational. Air-gap-capable, with role-based access scoped to care team, matter, or project.
The RAM-heavy single node. Flash-Next’s offloadable N-gram table lets us build nodes with modest accelerators and large system memory — a materially cheaper bill of materials for organizations whose workload is documents rather than code.
Questions we get
Frequently asked questions
What hardware does Qwen3.8-Flash-Next require?
Less than any other frontier-class model on the bench. The FP8 checkpoint is 172.78 GiB and BF16 335.28 GiB; Unsloth’s 4-bit GGUF is 111 GB and the 2-bit builds 72–79 GB, with the 51B N-gram table able to live in system RAM. A two-GPU workstation with 96 GB parts serves a quantized build to a workgroup; a four-GPU node (TP4 reference, TP2 minimum validated FP8 on GB300) serves FP8 firm-wide. Only 6B parameters activate per token, so decode is fast on all of them.
Is Qwen3.8 Apache-licensed?
It depends on the model. Qwen3.8-27B is Apache 2.0. Qwen3.8-Flash-Next ships under the Qwen Community License 1.0, and Qwen3.8-Max under its own custom license — both permit internal enterprise deployment in our reading, but the exact clauses should be reviewed by procurement before commercial redistribution. We include the license analysis in every deployment file, and where the license must be unambiguous we deploy Qwen3.8-27B or MIT-licensed GLM-5.3-Flash instead.
How accurate is open-weight document understanding?
Strong enough that we validate it the only way that matters: on your documents. Before rollout we benchmark the candidate model against a sample of your real records — handwriting, stamps, low-quality scans included — and report measured accuracy. The pipeline ships when it clears your bar, with human review kept in the loop where stakes demand it.
Why not just use a cloud OCR or vision API?
Because the documents are the sensitive data. Cloud OCR routes patient charts, privileged evidence, or controlled drawings through a third party’s infrastructure — with all the retention, logging, and jurisdiction exposure that entails. On-premise Qwen3.8 delivers the same capability with zero egress, and pairs it with reasoning the OCR services don’t have.
Make decades of paper searchable — without scanning it to a cloud.
The two-week sovereignty assessment sizes the hardware against your real workloads and hands you a written architecture with a cost model — before you buy a single GPU.
Book a sovereignty assessment