
MiMo-V2.5-Pro-UltraSpeed: a 1T model serving over 1000 tokens per second
Xiaomi and TileRT released MiMo-V2.5-Pro-UltraSpeed, which decodes over 1000 tokens per second on a trillion-parameter model. Access runs on application from June 9 to 23, at three times the price of MiMo-V2.5-Pro.
Xiaomi's MiMo team has released MiMo-V2.5-Pro-UltraSpeed, a serving configuration built with the systems company TileRT that, the team says, breaks 1000 tokens per second of decode speed on a trillion-parameter model for the first time.
Limited window and pricing
Access is application-based rather than open. The API is offered at a limited-time promotional price — three times the cost of MiMo-V2.5-Pro for roughly ten times the generation speed, which the company summarises as "3x the price, 10x the output experience". The offer runs from June 9 to June 23, 2026, 23:59 Beijing time (08:59 PDT), and is API-only: the Token Plan is not supported.
Applications go through platform.xiaomimimo.com/ultraspeed. Slots are limited, submission does not guarantee approval, and priority goes to enterprises and professional developers. Approved users also get free Chat access during the two-week window at ultraspeed.xiaomimimo.com, with each account allowed up to ten queue entries a day, sessions capped at 30 minutes, and idle sessions released after five minutes.
Where the speed comes from
The result is not a new architecture but a co-design effort between the model team and TileRT's inference system. Xiaomi notes that comparable extreme speeds in the industry usually depend on specialised hardware — Cerebras's wafer-scale integration or Groq's on-chip SRAM designs — while this one runs on commodity GPUs: 1000+ tokens per second from a 1T model on a single standard 8-GPU node.
Two model-side choices carry most of the weight. FP4 quantization (MXFP4) is applied only to the MoE experts, which hold the bulk of the parameters and tolerate quantization best, using quantization-aware training while the rest of the model keeps its original precision. DFlash speculative decoding then predicts a whole block of masked tokens in one forward pass instead of drafting autoregressively; its draft model uses sliding window attention, and block size is capped at eight to keep verification cheap.
Measured acceptance and open weights
The reported acceptance length — how many drafted tokens the large model confirms per verification round — is 6.30 in coding scenarios, peaking at 7.14, 5.56 in math and reasoning, and 4.29 in agent tasks. Acceptance is still low in open-ended conversation, the team says, and work there continues.
On the system side, TileRT contributes a persistent engine kernel that keeps the compute pipeline resident on the GPU instead of launching operators one by one, plus warp specialization that splits communication, data movement and tensor computation across thread groups. The resulting checkpoint, MiMo-V2.5-Pro-FP4-DFlash, is open-sourced on Hugging Face, and Xiaomi says UltraSpeed support for MiMo-V2.5 is planned next.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.