Back
Qwen3.8-Max: 2.4 Trillion Parameters, Open Weights and a New Bar for Coding
SiTech Team3 წთ. საკითხავი

Qwen3.8-Max: 2.4 Trillion Parameters, Open Weights and a New Bar for Coding

The Qwen team has released Qwen3.8-Max, its most capable model yet — 2.4 trillion parameters, already available via the QwenCloud API, with open weights promised for next week.

The Qwen team released Qwen3.8-Max on 3 August 2026 — the most capable model in the family to date, and the first Qwen-Max-class model whose weights will be open-sourced. Built on the architectural foundation of Qwen 3.5, it scales to 2.4 trillion parameters (95 billion of them active) and is available through the QwenCloud API; the open weights are promised for next week.

Coding: days of autonomous work

The team tested the model on three challenges in which results had to be earned by actually writing and running code, with no human help at all. First, it was asked to create the oh-my-cli project from scratch and spend over 10 days building a self-evolving harness that folds user feedback, community practices and the model’s own self-tests into one engineering loop. As of 30 July 2026, after roughly 16 days of fully autonomous operation, the repository had accumulated 265 commits, 127 pull requests and 151 issues.

Second, the model was handed a research paper — “Unified Data Selection for LLM Reasoning” — and asked to reproduce its experiment in code and then improve on it. In about five days (~125 hours) it wrote roughly 7,600 lines of code, took over 1,100 actions and ran 33 rounds of GPU training, reproducing the paper’s six main findings; the paper’s selection method beats picking data at random by 7.7% on AIME24. It then ran a self-improving research loop, testing 18 ideas of its own across four rounds, and evolved a method that beats the paper’s approach by 2.7 points on AIME24 (52.29%).

Third, it entered a live contest on Alibaba Cloud’s Tianchi platform, where 526 human teams were competing. Under a 24-hour limit, across 45 submissions, its accuracy climbed from 0.60 to 0.853, beating 458 of the 526 teams (87%).

Work and long-horizon tasks

The team also stresses that the model performs comparably across several agent harnesses (QwenWork, Claude Code, Codex, OpenClaw, Hermes), achieved by jointly scaling reinforcement-learning environments and compute. On benchmarks it scores 86.6 on Terminal Bench 2.1, 67.7 on SWE-bench Pro and 93.0 on PaperBench — the last above Opus 4.8 (80.3), Fable5 (88.8) and GPT-5.6 Sol (90.5). On long-horizon work, Qwen3.8-Max autonomously took a chip design block — a GCD/RSA cryptographic hardware accelerator — from 8,298 gates down to 678, and shrank the physical die from 106×106 µm² to 46×46 µm², an 81% reduction. In a 365-day e-commerce simulation that starts with ¥100,000 in capital, it ended with a balance of ¥416,252 — a 4.16x return, 38% ahead of the runner-up, GLM 5.2.

Availability

The model is already callable through the QwenCloud API, with a 1-million-token context window and up to 65,536 output tokens. Open weights are promised for next week — a first for a Qwen-Max-class model. The team also introduced Qwen-MM-Plugins, an extension library that adds image and video processing and other multimodal capabilities to existing agent harnesses.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.