Alibaba releases QwQ-32B-Preview reasoning model
Built on Qwen2.5-32B and released under an Apache 2.0 licence, Alibaba flagged the model could enter circular reasoning loops and mix languages mid-response.
- Open weights & ecosystem
- Models & capabilities
- Notable
Alibaba’s Qwen team released QwQ-32B-Preview, a 32.5-billion-parameter open-weight model built on top of Qwen2.5-32B and further trained to reason step by step before answering, explicitly positioned against OpenAI’s o1-preview. It was released under an Apache 2.0 licence on Hugging Face, making it usable and modifiable without the restrictions attached to o1, which remained closed and accessible only through OpenAI’s API.
Alibaba labelled the release an “experimental research model” and a preview rather than a finished product, and was unusually direct about its limitations in the accompanying documentation. The model was reported to be strong on mathematics and coding tasks but weaker on common-sense reasoning and nuanced language understanding than on formal problems, and Alibaba flagged that it could enter circular reasoning patterns that produced unusually long responses without reaching a conclusion, and could mix languages or switch unexpectedly mid-response. The Qwen team said the model required additional safety enhancement work before broader use.
QwQ-32B-Preview was among the earliest open-weight models to replicate the extended-reasoning, chain-of-thought approach that OpenAI’s o1 had introduced as a closed, proprietary system two months earlier — arriving weeks before DeepSeek’s R1 pushed the same idea further and drew far more attention. Its release demonstrated that the core technique behind test-time reasoning models was reproducible outside the lab that introduced it, at a fraction of the parameter count of typical frontier models, even as Alibaba itself described the specific model as unfinished.