
Qwen3.8 Max Tops Artificial Analysis's Agentic Index
Alibaba's flagship reasoning model has taken the top spot in the agentic index compiled by Artificial Analysis, a ranking built to measure how AI systems handle long-horizon work rather than exam-style questions.
Artificial Analysis has a new leader on its agentic index: Qwen3.8 Max, the flagship reasoning model that Alibaba released on August 3, 2026. The ranking, published on the independent evaluation site, places the model at the top of an index built specifically around agentic work — tasks in which a system has to plan, use tools and deliver a finished result rather than answer a question.
What the index measures
The agentic index draws on evaluations that Artificial Analysis runs in-house. Among them are AA-Briefcase v1.1, which tests long-horizon knowledge work through deliverables such as spreadsheets, presentations and memos; GDPval-AA v2.1, built on real-world work tasks; AutomationBench-AA, which covers SaaS workflows; and Terminal-Bench 4.0, aimed at coding and terminal use. Those agentic tests sit alongside classical knowledge benchmarks in the site's broader Intelligence Index, which currently combines ten evaluations.
The model behind the ranking
Qwen3.8 Max is a proprietary reasoning model, and Alibaba has not disclosed its parameter count. It accepts text, image and video input, returns text, and offers a context window of one million tokens. On the Artificial Analysis Intelligence Index it scores 40, well above the median of 24 for comparable models. The trade-offs are speed and verbosity: the model outputs roughly 41 tokens per second, at the slow end of its tier, and generated about 180 million tokens during the intelligence evaluation, twice the median. API pricing is $2 per million input tokens and $6 per million output tokens, with an 88 percent discount on cached input.
Why it matters
Agentic benchmarks have become the field's main proving ground because they approximate paid work instead of quizzes, and they are harder to satisfy by memorising answers. A proprietary model from Alibaba leading this ranking says as much about where the frontier of practical AI has moved as it does about a single release: competition now runs on whether models can complete multi-step jobs reliably, and on what that reliability costs.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.