
Qwen3.8-2.4T-A95B brings a Max-class model to open weights
The Qwen team has published Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter mixture-of-experts model that is the first Qwen-Max-class system to be released openly, with a 262K-token context window.
First Max-class model in open release
The Qwen team has published Qwen3.8-2.4T-A95B on Hugging Face, describing it as the most capable generation in the Qwen open-model family to date and the first time a Qwen-Max-class model has been released openly. The weights ship in the Transformers format and are compatible with serving stacks such as vLLM, SGLang and TokenSpeed; the model card lists a qwen3.8-max licence and reports 55,519 downloads in the past month.
Architecturally it is a mixture-of-experts model: 2.4 trillion parameters in total and 95 billion activated, spread over 92 layers with 512 experts, of which ten routed experts plus one shared expert fire per token. The stack alternates Gated DeltaNet and Gated Attention blocks, and the model was trained with multi-token prediction. Context length is 262,144 tokens natively, extensible up to 1,010,000.
Benchmark results
The published comparison puts Qwen3.8-Max — the hosted version built on these weights — against Opus 4.8, Fable 5 and GPT-5.6 Sol (max). It scores 86.6 on Terminal Bench 2.1, 67.7 on SWE-bench Pro, 56.6 on DeepSWE 1.1 and 93.0 on PaperBench. On general capabilities it reports 92.6 on GPQA Diamond, 43.6 on HLE and 82.8 on IFBench, plus 92.9 on the 256K-token MRCR v2 retrieval test.
The lead is not across the board. Fable 5 stays ahead on SWE-bench Pro, 80.0 to 67.7, and GPT-5.6 Sol (max) posts 88.8 on Terminal Bench 2.1 against Qwen's 86.6.
Text only, and thinking cannot be switched off
The released checkpoint is text-only and requires thinking mode for every interaction: multimodal input is not supported and reasoning cannot be disabled, so each answer begins with reasoning wrapped in think tags before the final output. Reasoning depth can be tuned through a reasoning_effort parameter, while preserve_thinking keeps reasoning context from earlier messages and is on by default for all workloads. For agentic work the card recommends allocating up to 262,144 output tokens for reasoning and 131,072 for the final response.
The Max branding applies to the hosted build: the team says Qwen3.8-Max adds vision input, a non-thinking mode, 1M context by default and official built-in tools on top of the open weights.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.