Model bench · Alibaba

Deploy Qwen3.8 on-premise.

The Qwen3.8 family, released through August 2026, is the efficiency corner of the open-weight bench. Qwen3.8-Flash-Next (August 26) is a 125-billion-parameter mixture-of-experts with just 6B active parameters, a 51B N-gram embedding table designed to live in system RAM rather than VRAM, a 4B multi-token-prediction head, native text-image-video input, and 262K context extensible to 1M — an early preview of the Qwen4 architecture. Qwen3.8-27B is the dense, Apache-2.0, multimodal workstation model. Both read what your organization actually holds — scanned records, drawings, forms, evidence — without a page leaving your network.

Specifications

Qwen3.8 at a glance.

SpecQwen3.8
Qwen3.8-Flash-Next125B MoE · 6B active · 512 experts (10 routed + 1 shared) · 51B N-gram table + 4B MTP head · hybrid Gated DeltaNet + Qwen Sparse Attention · 48 layers
Flash-Next modalities & contextText, image, video in; text out · 262,144 native → 1,000,000 with YaRN
Flash-Next sizesFP8 172.78 GiB · BF16 335.28 GiB · Unsloth GGUF: UD-Q4_K_XL 111 GB · UD-Q2_K_XL 78.9 GB · UD-IQ1_S 72.5 GB · Qwen/Qwen3.8-Flash-Next
Flash-Next benchmarks (card)SWE-bench Pro 62.5 · DeepSWE 58.7 · GPQA Diamond 91.7 · LiveCodeBench v6 91.9 · CoWorkBench 73.9 · MathVision 95.7
Flash-Next licenseQwen Community License 1.0 — not Apache 2.0; procurement should read the terms before commercial redistribution
Qwen3.8-27B27B dense · multimodal · 262K → 1M · Apache 2.0 · SWE-bench Pro 61.7 · Terminal-Bench 2.1 73.0 · Qwen/Qwen3.8-27B
Qwen3.8-Max (2.4T / 95B active)Open weights under a custom license (Qwen/Qwen3.8-2.4T-A95B) · text-only · 1M context — cluster-tier only
ServingvLLM · SGLang · TokenSpeed · KTransformers · llama.cpp · Transformers; TP4 in the reference configs, TP2 minimum validated FP8 on GB300

The N-gram table is the genuinely new idea: 51B parameters of bigram/trigram embeddings that “can be offloaded to host memory and overlapped with model computation.” That moves a large share of the model into cheap system RAM — a new axis for single-node sizing that no other frontier-class model currently offers.

Hardware requirements

Sizing the deployment honestly.

Capabilities

What it’s best at.

Qwen3.8 earns its place on regulated benches because the most sensitive material is so often paper and pixels: patient charts, discovery productions, engineering-change orders, handwritten notes, screen recordings. Flash-Next reads documents, images and video natively, extracts structure, and reasons over it in one pass — then hands clean output to whatever language model sits beside it. Its CoWorkBench (73.9) and AndroidWorld (84.5) scores make it a strong agent runtime for tool-driven back-office work.

Against GLM-5.3-Flash — the other efficiency flagship of late August — Flash-Next is smaller and cheaper to run but carries the Qwen Community License rather than MIT and a 262K native context. Where the license must be permissive with no reading required, Qwen3.8-27B (Apache 2.0) or GLM-5.3-Flash (MIT) are the choices; where every gigabyte of VRAM matters, Flash-Next is. We benchmark on your documents before choosing.

Deployment patterns

How we deploy it.

The intake pipeline. Documents enter, structured data leaves — inside your network. Flash-Next performs understanding and extraction, a language model classifies and routes, and nothing transits a third party at any step.

Multimodal retrieval. Paired with GLM-5.3 or GLM-5.3-Flash over a PLM, DMS, or records archive, Qwen3.8 indexes what scanners captured and text search never could — making decades of paper conversational. Air-gap-capable, with role-based access scoped to care team, matter, or project.

The RAM-heavy single node. Flash-Next’s offloadable N-gram table lets us build nodes with modest accelerators and large system memory — a materially cheaper bill of materials for organizations whose workload is documents rather than code.

Questions we get

Frequently asked questions

What hardware does Qwen3.8-Flash-Next require?

Less than any other frontier-class model on the bench. The FP8 checkpoint is 172.78 GiB and BF16 335.28 GiB; Unsloth’s 4-bit GGUF is 111 GB and the 2-bit builds 72–79 GB, with the 51B N-gram table able to live in system RAM. A two-GPU workstation with 96 GB parts serves a quantized build to a workgroup; a four-GPU node (TP4 reference, TP2 minimum validated FP8 on GB300) serves FP8 firm-wide. Only 6B parameters activate per token, so decode is fast on all of them.

Is Qwen3.8 Apache-licensed?

It depends on the model. Qwen3.8-27B is Apache 2.0. Qwen3.8-Flash-Next ships under the Qwen Community License 1.0, and Qwen3.8-Max under its own custom license — both permit internal enterprise deployment in our reading, but the exact clauses should be reviewed by procurement before commercial redistribution. We include the license analysis in every deployment file, and where the license must be unambiguous we deploy Qwen3.8-27B or MIT-licensed GLM-5.3-Flash instead.

How accurate is open-weight document understanding?

Strong enough that we validate it the only way that matters: on your documents. Before rollout we benchmark the candidate model against a sample of your real records — handwriting, stamps, low-quality scans included — and report measured accuracy. The pipeline ships when it clears your bar, with human review kept in the loop where stakes demand it.

Why not just use a cloud OCR or vision API?

Because the documents are the sensitive data. Cloud OCR routes patient charts, privileged evidence, or controlled drawings through a third party’s infrastructure — with all the retention, logging, and jurisdiction exposure that entails. On-premise Qwen3.8 delivers the same capability with zero egress, and pairs it with reasoning the OCR services don’t have.

Make decades of paper searchable — without scanning it to a cloud.

The two-week sovereignty assessment sizes the hardware against your real workloads and hands you a written architecture with a cost model — before you buy a single GPU.

Book a sovereignty assessment