MiniMax releases MiniMax-M3, combining frontier coding, 1M context and native multimodality
The 428-billion-parameter model (23bn active) reached a 1M-token context window and, MiniMax said, outscored GPT-5.5 and Gemini 3.1 Pro on SWE-Bench Pro at a fraction of the price.
- Open weights & ecosystem
- Models & capabilities
- Benchmarks & progress
- Notable
MiniMax released MiniMax-M3, an open-weight mixture-of-experts model with 428 billion total parameters and roughly 23 billion active per token, extending its context window to 1 million tokens — five times that of its predecessor, M2.7. The model takes text, image and video input natively rather than through bolt-on modules, and MiniMax attributed its throughput to a custom attention mechanism, MiniMax Sparse Attention, which it said cuts per-token compute at long context to about a twentieth of the previous generation’s.
On SWE-Bench Pro, MiniMax reported a score of 59.0%, ahead of the figures the company cited for GPT-5.5 and Gemini 3.1 Pro, though behind Claude Opus 4.7. On BrowseComp, a web-browsing agent benchmark, MiniMax put its score at 83.5 against 79.3 for Opus 4.7. Independent replication of these comparisons was not available at release. Pricing was set at roughly $0.30 per million input tokens, more than fifteen times cheaper than Anthropic’s listed rate for Opus-class access — continuing a pattern in which Chinese labs have paired near-frontier coding performance with aggressively low API prices. Weights were published on Hugging Face under a custom MiniMax community licence rather than a fully permissive one.
The release added to a crowded month for open-weight coding models, arriving alongside comparable pushes from Moonshot AI, Zhipu and others, and reinforced the argument that the gap between the best closed and open models on agentic coding tasks had narrowed to a matter of months rather than generations.