
Kimi K2.7 Code: Moonshot’s Open Coding Model Cuts Thinking Tokens by ~30%
Moonshot AI’s Kimi K2.7 Code builds on K2.6 with a trillion-parameter MoE design, a 256K context window and roughly 30% fewer thinking tokens — but its own benchmark table still shows GPT-5.5 and Claude Opus 4.8 ahead on most tests.
Moonshot AI has published Kimi K2.7 Code, a coding-focused agentic model built on Kimi K2.6. The weights and the code repository are released under a Modified MIT licence, and the model card frames the release around two claims: stronger end-to-end completion of long, real-world software engineering tasks, and a reduction of roughly 30% in thinking-token usage compared with its predecessor.
Same architecture, more work per token
The specification is unchanged from the K2.5/K2.6 line: a mixture-of-experts transformer with 1 trillion total parameters, 32 billion activated per token, 384 experts of which 8 are selected plus one shared expert, 61 layers (one dense), and a 256K-token context window. Attention uses MLA and the activation function is SwiGLU; a 400-million-parameter MoonViT encoder adds image and video input, though video remains experimental and limited to the official API. Because the architecture matches earlier versions, Moonshot says existing deployments can be reused, and the model keeps native INT4 quantization.
What the benchmark table shows
On the company’s in-house Kimi Code Bench v2, K2.7 Code scores 62.0 against 50.9 for K2.6. It also improves on Program Bench, where an agent must rebuild a program’s behaviour from a compiled binary and documentation alone, from 48.3 to 53.6, and on MLS Bench Lite from 26.7 to 35.1. Agentic results move in the same direction: 46.9 on the long-horizon Kimi Claw 24/7 Bench (up from 42.9), 76.0 on MCP Atlas (from 69.4) and 81.1 on MCPMark-Verified (from 72.8).
The same table keeps the comparisons honest. GPT-5.5 scores 69.0, 69.1 and 35.5 on the three coding benchmarks, while Claude Opus 4.8 reaches 67.4, 63.8 and 42.8; on the agentic sets the two lead at 52.8/79.4/92.9 and 50.4/81.3/76.4 respectively. K2.7 Code tops none of the six rows outright, but closes part of the distance to both closed models.
Availability
Moonshot recommends running the model on vLLM, SGLang or KTransformers, with transformers between 4.57.1 and 5.0.0, and offers an OpenAI- and Anthropic-compatible API through platform.moonshot.ai. The model forces thinking mode and preserve_thinking, which retains reasoning content across multi-turn interactions. The Hugging Face page recorded about 175,000 downloads in the last month and 45 Spaces built on the model.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.