Back
Kimi K3: an open model with 2.8 trillion parameters
SiTech AI Team3 წთ. საკითხავი

Kimi K3: an open model with 2.8 trillion parameters

Moonshot AI's Kimi K3 is a 2.8-trillion-parameter model with native vision and a one-million-token context window. Weights are due by July 27, 2026, and the company says it still trails the top proprietary models overall.

Moonshot AI's Kimi has introduced Kimi K3, its most capable model so far: a 2.8-trillion-parameter system with native vision and a one-million-token context window. The company describes it as the first open model in the three-trillion-parameter class.

Architecture and scaling

K3 is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), two changes aimed at how information flows across sequence length and model depth. Its mixture-of-experts layer activates 16 of 896 experts under a Stable LatentMoE framework. With revised training and data recipes, Kimi says the result is roughly a 2.5x gain in scaling efficiency over Kimi K2. For nine of the past twelve months, Kimi models have set the upper bound of open-model sizes.

Long-horizon coding

In kernel optimization tests, each model worked independently in an identical sandbox for up to 24 hours on four tasks spanning AttnRes, KDA and a 512-head-dimension MLA kernel across NVIDIA Hopper GPUs and general-purpose GPUs from an alternative vendor. Kimi reports that K3 performed competitively with Claude Fable 5, which was evaluated by a third party and may include fallback behaviour, and clearly ahead of Claude Opus 4.8, GPT 5.6 Sol and GPT 5.5.

K3 also built MiniTriton, a compact Triton-like compiler with its own tile-level IR layer over MLIR and a PTX code-generation pipeline. On supported roofline benchmarks it matched or beat Triton and torch.compile on some workloads and sustained end-to-end nanoGPT training with stable convergence. In another autonomous run, an early K3 built and verified a chip for a nano model based on its own architecture in 48 hours with open-source EDA tools on the Nangate 45nm library: 4 mm², timing closed at 100 MHz and over 8,700 tokens/s decode throughput in simulation.

Knowledge work, and the gaps

On research tasks, K3 reproduced the I–Love–Q universal relations in computational astrophysics by cross-checking more than 20 papers, evaluating over 300 equations of state and generating 3,000+ lines of Python — work Kimi says would normally take an experienced researcher one to two weeks, done in about two hours. Internal case studies include a 42-year history of the AI ASIC industry and an analysis of 391 gravitational-wave events run with more than 20 concurrent subagents.

K3 is available now on Kimi.ai, Kimi Work, Kimi Code and the Kimi API, running at maximum thinking effort by default. Full model weights are scheduled for release by July 27, 2026, with a technical report to follow. Kimi is candid about the limits: K3 still trails Claude Fable 5 and GPT 5.6 Sol overall, is sensitive to missing thinking history in an agent harness, and can act on its own assumptions when user intent is ambiguous.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.