Mistral 7B beats larger models and ships by torrent
A magnet link with no blog post or announcement stood in deliberate contrast to the polished launches of Meta and Google, and the 7-billion-parameter model still beat Llama 2 13B.
- Open weights & ecosystem
- Models & capabilities
- Major
Mistral AI, a French startup founded a few months earlier, released its first model — a 7.3-billion-parameter language model — by posting a bare magnet link on social media, with no blog post, press release or launch event to accompany it. A short page on Mistral’s own site followed, but the model itself reached the world first as a torrent.
Mistral claimed the model outperformed Meta’s Llama 2 13B on every benchmark it reported, and Llama 1 34B on many, despite having roughly half or a fifth as many parameters — a claim read at the time as evidence that architecture and training choices, not just scale, still drove meaningful gains. The model used grouped-query attention for faster inference and sliding-window attention to handle longer sequences with less memory, and was released under the permissive Apache 2.0 licence, allowing unrestricted commercial use.
The release method was itself a statement. Coming weeks after Google’s more conventional, marketing-heavy model launches, and in a period when even nominally “open” releases from larger labs came wrapped in blog posts, licensing caveats and usage restrictions, Mistral’s magnet-link-only release was read by much of the AI community as a pointed contrast: open-weight AI stripped back to the weights themselves, distributed the way pirated media had been for two decades, with no corporate ceremony attached.
The release established Mistral’s reputation as a serious, well-resourced new entrant capable of competing with Meta’s open-weight models on a fraction of the parameter count, and set the tone — technically strong, deliberately minimal in presentation — for the company’s subsequent releases as it built toward the larger closed and open models that followed.