Mistral launches Mistral 3 model family
Mistral released Mistral 3, including dense models at 3B/8B/14B and a new mixture-of-experts Mistral Large 3 (41B active, 675B total), all under Apache 2.0.
- Models & capabilities
- Open weights & ecosystem
- Notable
Mistral AI released the Mistral 3 family: three dense “Ministral 3” models at 3B, 8B and 14B parameters, each offered in base, instruct and reasoning variants, and Mistral Large 3, a new sparse mixture-of-experts model with 41 billion active and 675 billion total parameters — Mistral’s first MoE release since Mixtral. All were published under the Apache 2.0 licence.
Mistral said Large 3 was trained from scratch on 3,000 Nvidia H200 GPUs and ranked second among open-weight non-reasoning models on the LMArena leaderboard, with performance roughly matching the strongest open-weight instruction-tuned models on general tasks and particular strength in multilingual conversation. The Ministral models add native image understanding and support for more than 40 languages; Mistral claimed the 14B reasoning variant reached 85% on the AIME ’25 maths competition benchmark, and said the family delivered the best cost-to-performance ratio of any open-weight model it compared against, using markedly fewer output tokens than competitors for comparable tasks. These figures were Mistral’s own, without independent verification at release.
The release kept Mistral in its position as the most prominent Western lab still shipping frontier-adjacent open weights under a permissive licence, at a time when OpenAI, Anthropic and Google DeepMind kept their flagship models closed. Models were made available immediately through Mistral AI Studio, AWS Bedrock, Azure and Hugging Face, with NVIDIA NIM and SageMaker support to follow.