
Ai2 releases Olmo-core 3, an open training system for trillion-parameter MoEs
Ai2 has released Olmo-core 3, a redesigned open training system for mixture-of-experts models that scales into the trillion-parameter range. In one benchmark the expert pool grew from 8 to 128 experts while throughput fell by less than 5%.
Ai2 (Allen Institute for AI) released Olmo-core 3 on October 1, a redesign of its open framework built around a new training system for mixture-of-experts (MoE) models. According to Ai2, it scales MoE training into the trillion-parameter range while preserving efficiency and is one of the core systems behind the next generation of Olmo.
MoE models use only part of the model for each input, but storing the full model and routing data to the right experts adds costs that grow with size, eroding that advantage.
A new training stack for MoEs
In one benchmark, the expert pool grew from 8 to 128 while only four experts were selected per token, keeping active parameters per token near 3.2B. Total parameter capacity grew from 4.6B to 47B, and training throughput fell by less than 5%.
The earlier implementation used fully sharded data parallelism (FSDP), regathering weights for every batch. Olmo-core 3 moves to a distributed data parallelism (DDP) design: experts stay resident on GPUs and data is routed to them.
On eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, versus 19,400 with the earlier implementation, about 2.7x the throughput.
Scaling into the trillion-parameter range
Olmo-core 3 splits models across hardware with expert parallelism, pipeline parallelism and a distributed optimizer. Routing metadata stays on the GPUs, routed data lands directly in expert input buffers, and grouped GEMM merges many small expert computations.
The stack also supports MXFP8, a lower-precision format: on four B300 GPUs it delivered about 21% higher throughput than BF16, with peak active memory down from 103 GiB to 95 GiB.
Ai2 also benchmarked a 1.2-trillion-parameter model with 58.36 billion parameters active per token across 512 GPUs, reaching up to 858 TFLOP/s per GPU; the tests used random routing to measure system performance, not model quality. A DeepEP v2 experiment reached 2.38 trillion total parameters in a short-capacity test.
Built for the next Olmo, open for everyone
The next-generation Olmo will use an MoE architecture and should be the most capable yet, trained on the largest dataset with the longest context window. The stack is fully open: researchers can train their own MoEs and adapt it to different hardware; the technical report, code and an interactive demo are available now.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.