Moonshot AI releases Kimi K2, a 1-trillion-parameter open-weight model
The mixture-of-experts model activates 32 billion of its 1 trillion parameters per token and was trained with the Muon optimiser at a scale its makers said had previously caused instability.
- Open weights & ecosystem
- Models & capabilities
- Major
Moonshot AI released Kimi K2, a mixture-of-experts language model with 1 trillion total parameters, of which 32 billion are active for any given token across 384 experts plus one shared expert. The company published weights for both a base model, intended for further fine-tuning, and an instruction-tuned variant built for chat and agentic use, under a modified MIT licence that made the weights freely usable, including commercially.
Moonshot said it pre-trained the model on 15.5 trillion tokens using the Muon optimiser, a technique it described as prone to instability at this scale until the company developed fixes to keep training stable. On released benchmarks the model reported strong coding and tool-use results, including a majority pass rate on LiveCodeBench and high scores on agentic tool-use evaluations, putting it ahead of several then-current closed models such as GPT-4.1 and Claude Sonnet 4 on coding-specific tests, according to Moonshot’s own figures.
Commentary compared the release to DeepSeek’s earlier disruption of the closed-model market, on the grounds that a lab outside the usual frontier-lab set had again published a model with near-frontier agentic and coding capability as open weights rather than as an API-only product. The release also marked a strategic shift for Moonshot, which had previously been known chiefly for long-context products, toward positioning itself as a lab pushing the frontier of open-weight capability specifically in coding and agentic tasks — an area where open models had lagged closed ones more than they had on general knowledge benchmarks.