
Inception's Mercury 2.5 diffusion LLM hits 770 tokens per second
Inception's diffusion-based Mercury 2.5 reaches 770.4 output tokens per second in Artificial Analysis testing, with a 260k context window and $0.25 per million input tokens at launch.
Inception released Mercury 2.5 on September 8, 2026 — a proprietary diffusion large language model (dLLM) that the company describes as the fastest reasoning LLM in production. Unlike autoregressive models that write text token by token, Mercury 2.5 builds its answer in parallel and refines it in stages.
770 tokens per second
Artificial Analysis measured the model at 770.4 output tokens per second through Inception's API — far above the median of 110.3 tokens per second for reasoning models in its price tier. Inception itself claims more than 1,100 tokens per second in production. Time to first answer token is 2.91 seconds, slightly above the tier median of 2.22 seconds.
Mercury 2.5 accepts and outputs text only: it cannot process images and is not multimodal. The model is proprietary, and its weights are not publicly available. Its context window is 260,000 tokens, roughly 390 A4 pages, up from 128,000 in Mercury 2.
Benchmarks and price
On the Artificial Analysis Intelligence Index v4.3.2, Mercury 2.5 scores 12, matching the median for comparable models. It used 35 million output tokens on the index — less than half the median of 84 million, which Artificial Analysis calls fairly concise. Evaluating one Intelligence Index task costs $0.06 on average.
Pricing through Inception's API is $0.25 per million input tokens and $0.75 per million output tokens, with a 90% cache discount and a blended rate of $0.14 per million tokens. At launch, an 80% discount on OpenRouter brings the price to $0.04 and $0.15 per million tokens.
Why it matters
Diffusion generation targets low-latency workloads — search, voice assistants and coding agents — where speed and cost matter as much as raw capability. Inception says Mercury 2.5 gains 10 points of intelligence over Mercury 2 and is comparable to cost-optimized frontier models such as GPT-5.6 Luna, Gemini 3.5 Flash-Lite and Claude Haiku 4.5. The model adds tunable reasoning, native tool use, JSON mode and parallel tool calls, and is available via Inception's API, OpenRouter and Baseten. Inception says it has started training its largest model yet and plans to release it in the coming months.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.