Qwen3.8-2.4T-A95B
modelYour notes
Alibaba's first open-weight Qwen-Max-class release and its largest public model: a 2.4T-parameter MoE with 95B active per token. Its 92-layer hybrid stack interleaves three Gated DeltaNet linear-attention blocks with one gated-attention block; 10 of 512 routed experts plus one shared expert are active. Native context is 262K and can extend to 1.01M tokens.
Built for coding, professional work, research, and long-horizon agents, with controllable reasoning effort and preserved multi-turn thinking. The weights use the custom Qwen3.8-Max license: modification and redistribution are allowed, but large consumer products must display attribution and qualifying MaaS/work- assistant businesses must obtain a separate commercial license.
Qwen3.8-Max is the same model. The API flagship announced August 3, 2026 (GA successor to the July WAIC preview; blog title "Qwen3.8-Max: A New Bar for Coding and Cowork") is, per the HF card, "the official version based on Qwen3.8-2.4T-A95B" with vision input, a non-thinking mode, 1M context by default and built-in tools, at $2/$6 per M tokens; the open weights followed on August 12. Artificial Analysis scores the two endpoints separately (40 for both on v4.3; the top-level score here anchors to the open-weight page). The generation also includes the dense Qwen3.8-27B (August 14, Apache 2.0, native vision, Qwen3.5-style GDN hybrid) and the Qwen3.8-Flash-Next architecture preview.
Model Details
Benchmark Scores
| Benchmark | Score | Mode |
|---|---|---|
| Terminal-Bench 2.1 | 86.6% | — |
| SWE-bench Pro | 67.7% | — |
| GPQA Diamond | 92.6% | — |
Variants
| Name | Parameters | Notes |
|---|---|---|
| Qwen3.8-2.4T-A95B | 2.4T | Open weights, released Aug 12, 2026 (BF16 and FP8 repos); text-only, 92 layers, 512 experts (10 routed + 1 shared), Qwen3.8-Max license. AA Intelligence Index v4.3 score 40 (file's top-level anchor). |
| Qwen3.8-Max-0902 | — | AA Intelligence Index v4.3: 45 (the hosted Max page; the file anchors on the open 2.4T-A95B weights at 40). Sep 2 2026 snapshot of the hosted Max: deeper coding, multi-tool agent orchestration, chart and document vision; same 2.4T MoE, 1M context, unchanged $2/$6 pricing; the qwen3.8-max endpoint auto-routes to it from Sep 5 |
| Qwen3.8-Max | — | QwenCloud API version of the same 2.4T-A95B weights (announced Aug 3, 2026): adds vision input, non-thinking mode, 1M default context, built-in tools; $2/$6 per M tokens. AA scores it as a separate endpoint (qwen3-8-max): Intelligence Index v4.3 score 40. |
| Qwen3.8-27B | 27B | AA Intelligence Index v4.3: 34 (reasoning, xhigh; 28 medium, 26 low, 22 non-reasoning). Dense VLM sibling released Aug 14, 2026 under Apache 2.0: 64 layers (48 Gated DeltaNet + 16 gated attention), hidden 5120, 262K native / 1M extended context, image and video input. SWE-bench Pro 61.7, Terminal-Bench 2.1 73.0, GPQA Diamond 89.2. AA Intelligence Index v4.1.1 score 52 at xhigh effort (44 medium, 43 low, 35 non-reasoning). |