
Mistral Large 4 Ranked Best AI Model Outside U.S. and China
Mistral's new trillion-parameter Mistral Large 4 scores 38 on the Artificial Analysis Intelligence Index, making it the most intelligent model from outside the U.S. and China, though Chinese open models still lead overall.
Launch and Intelligence Index Ranking
Mistral launched its trillion-parameter Mistral Large 4 (ML4) model this week in Research Public Preview, saying it pushes the frontier of open-weight performance. The mixture-of-experts model, with 1.05 trillion total parameters and 49B to 52B active, scored 38 on the Artificial Analysis Intelligence Index v4.3.2, a result the independent benchmarking firm says makes it the most intelligent model from outside the U.S. and China.
The score is comparable to GPT-6 Luna (max, 38) and DeepSeek V4.1 Flash (max, 39), and places ML4 ahead of models from countries such as South Korea and the United Arab Emirates. However, at least five Chinese open models scored higher: Xiaomi's MiMo-V2.6-Pro at 46, Z.ai's GLM-5.3 (max) at 45, Moonshot's Kimi K3 (max) at 44, GLM-5.3-Flash at 42, and DeepSeek V4.1 Flash (max) at 39.
Cybersecurity Strengths
ML4 scored 50 on the Artificial Analysis Cyber Index, level with GLM-5.3-Flash and behind MiMo-V2.6-Pro at 56. Its strongest result came on CyberGym-E2E-AA, where it scored 82%, ahead of MiMo-V2.6-Pro (79%) and GPT-6 Luna (max, 78%). Mistral says the model solves 93% of Cybench, a set of 40 challenges drawn from security competitions, and CEO Arthur Mensch said in Abu Dhabi that the model is above the Chinese models on certain aspects, including cyber, according to Reuters.
The cyber result is only one of three tests in the Cyber Index. ML4 scored 16% on DeepsecBench-AA and 51% on CWE-Bench-AA. Artificial Analysis projects that once the weights ship, the model will rank among the top three open weights models on the Cyber Index. Mistral positions ML4 as state of the art among open models on enterprise workloads such as cybersecurity, finance and law, and says its 1.6B-parameter vision encoder lets it surpass even frontier closed models on visual grounding. Other reported results include DeepSWE v1.1 at 61.7%, Coding Agent Index at 49.8%, AutomationBench at 59.9% and Dense 200 at 42%.
Cost and Speed
Cost remains a challenge. ML4 comes in at $1.13 per AA task at list price, or $0.57 at the 50% discounted launch price, compared with $0.13 for MiMo-V2.6-Pro and $0.25 for GLM-5.3-Flash. Standard pricing is $1.36/$4.18 per 1M input/output tokens, with $0.14 per 1M cached input tokens. The higher cost stems from needing 200 million output tokens to run AA's index, against a median of 81 million. On speed, ML4 generated 116.1 tokens per second with 1.46 seconds to the first token, about 33% faster than its tier median and 2.6x faster to first token, as measured on Mistral's own API.
Training and Availability
ML4 was trained from scratch on 3,800 Nvidia Grace Blackwell GPUs in Mistral's own datacenters in Europe over roughly two months, up from 3,000 Nvidia H200s used for the smaller Large 3, which had 675B total parameters and 41B active. The model supports a 512k token context window, text and image input with text output, and its API now accepts 100 images per request, up from 8 for previous Mistral models. Mistral raised a €3 billion Series D, which it calls the largest equity round ever raised by a European technology company, and Mistral Compute has planned for 18,000 Grace Blackwell chips. Open weights, architecture details and additional benchmarks are expected before the end of October, at which point the model may be re-scored.
Sources: Tomshardware
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.