Timeline

Mistral AI releases Mistral Large 2

The 123-billion-parameter model reported 84.0% on MMLU and was released under a non-commercial research licence, with a separate paid licence for commercial use.

  • Models & capabilities
  • Notable

Mistral AI released Mistral Large 2, a 123-billion-parameter model with a 128,000-token context window, positioning it against GPT-4o, Claude 3 Opus and Meta’s Llama 3.1 405B on code and reasoning tasks. The company reported 84.0% accuracy on MMLU for the pretrained model and said the release cut hallucination rates by training the model to acknowledge when it lacked sufficient information to answer confidently, rather than guessing.

The model supported more than 80 programming languages and, Mistral said, performed well across a dozen natural languages including French, German, Spanish, Chinese, Japanese, Korean, Arabic and Hindi — a deliberate emphasis on multilingual capability that distinguished it from competitors trained predominantly on English data. It also added support for parallel and sequential function calling, aimed at developers building tool-using applications rather than simple chat interfaces.

Mistral released the weights under its own Mistral Research License, permitting free use for research and non-commercial purposes but requiring companies to negotiate a separate commercial licence to deploy the model in production — a middle position between fully open licences such as Apache 2.0 and the closed-weight approach of OpenAI and Anthropic. This licensing structure reflected an argument Mistral had made since its founding: that publishing weights did not have to mean forgoing revenue from the companies most likely to profit from them.

The release came a week after Meta’s Llama 3.1 405B and roughly six weeks before OpenAI’s o1 preview shifted competitive attention toward reasoning models, placing Mistral Large 2 among the last releases of the year judged primarily on conventional benchmark parity with GPT-4-class models rather than on chain-of-thought reasoning performance.