Moonshot AI releases Kimi K2.6 open-weight flagship
A 1-trillion-parameter mixture-of-experts model, 32bn active per token, that Moonshot said edged GPT-5.4 on SWE-Bench Pro while costing several times less to run.
- Open weights & ecosystem
- Models & capabilities
- Benchmarks & progress
- Notable
Moonshot AI released Kimi K2.6, an open-weight update to its Kimi K2 line, keeping the same overall architecture — a 1-trillion-parameter mixture-of-experts model with 32 billion parameters active per token, 384 experts, and multi-head latent attention — while extending context to 256,000 tokens and adding a 400-million-parameter vision encoder for native image and video input. Weights were published on Hugging Face under a Modified MIT Licence, with a branding requirement for products that exceed 100 million monthly active users or $20 million in monthly revenue.
On SWE-Bench Pro, Moonshot reported a score of 58.6, ahead of the 57.7 it credited to OpenAI’s GPT-5.4 on its highest reasoning setting, and ahead of the Claude Opus 4.6 and Gemini 3.1 Pro scores it cited for comparison. On Humanity’s Last Exam with tool use, it reported 54.0%, again the highest of the models compared. Moonshot said the model cost roughly five times less per input token and more than seven times less per output token than GPT-5.5 to run. These are vendor-reported figures from Moonshot’s own comparison tables, not independently verified benchmark runs.
The release also expanded Agent Swarm, Moonshot’s orchestration layer for running many sub-agents on one task, from 100 sub-agents and 1,500 coordinated steps in K2.5 to 300 sub-agents and 4,000 steps, with demonstrations of runs lasting over 13 hours and making more than 1,000 tool calls. As with the benchmark scores, these demonstrations came from Moonshot rather than third-party testing.
Kimi K2.6 continued a run of Chinese open-weight releases claiming parity with, or advantages over, the leading closed US models on specific coding and agentic benchmarks, at a fraction of the inference cost, while trailing on general reasoning tests such as AIME and GPQA Diamond.