Timeline

Alibaba releases Qwen3 model family

Open-weight family (dense and MoE, up to 235B-A22B) trained on 36 trillion tokens across 119 languages, Apache 2.0.

  • Open weights & ecosystem
  • Models & capabilities
  • Major

Alibaba released Qwen3, its third-generation open-weight model family, as eight models spanning six conventional dense architectures — from 0.6 billion to 32 billion parameters — and two mixture-of-experts models that activate only a fraction of their total parameters per token: a 30-billion-parameter model using 3 billion active parameters, and a 235-billion-parameter model using 22 billion active parameters. All eight were released under the permissive Apache 2.0 licence, allowing commercial use and modification without the restrictions some rival open releases carried.

The models were pretrained on approximately 36 trillion tokens spanning 119 languages and dialects, roughly double the 18 trillion tokens used for the prior Qwen2.5 generation, drawing on web text, PDF documents and synthetic data generated by Alibaba’s own maths and coding models. Alibaba said the largest model, Qwen3-235B-A22B, produced results comparable to DeepSeek-R1, OpenAI’s o1 and Gemini 2.5 Pro despite activating far fewer parameters at inference time than a comparably capable dense model would need, and that the smaller 30-billion-parameter mixture-of-experts model outperformed the earlier QwQ-32B model while using a tenth of its active parameters.

Like other Qwen releases, Qwen3 followed the pattern set by DeepSeek’s R1 release earlier in 2025: matching or approaching frontier closed-model performance on public benchmarks while publishing weights openly, at a fraction of the parameter count. The release consolidated Alibaba’s position, alongside DeepSeek, as one of the most prolific sources of competitive open-weight models over the course of 2025, reinforcing arguments from open-weight advocates that closed labs’ capability lead over openly available models was narrowing rather than widening.