Timeline

Alibaba releases QwQ-32B (full release)

Alibaba's Qwen team said reinforcement learning let a 32-billion-parameter model reach performance comparable to DeepSeek-R1's 671-billion-parameter model, under an Apache 2.0 licence.

  • Open weights & ecosystem
  • Models & capabilities
  • Benchmarks & progress
  • Notable

Alibaba’s Qwen team released the full version of QwQ-32B, a 32-billion-parameter reasoning model, under the permissive Apache 2.0 licence — allowing unrestricted commercial use and modification of the weights. A preview version had circulated the previous November.

Alibaba’s central claim was one of efficiency rather than raw capability: QwQ-32B, the company said, achieved performance “comparable to DeepSeek-R1,” a model with 671 billion total parameters (37 billion active per token) — roughly twenty times QwQ’s size. The company attributed the result to reinforcement learning applied to a strong pretrained base model, arguing that scaling RL, rather than scaling parameter count, was now driving gains on mathematical reasoning, coding and general problem-solving benchmarks. Evaluations in the release compared QwQ-32B against DeepSeek-R1’s smaller distilled variants and OpenAI’s o1-mini as well as against R1 itself.

The release arrived amid an active period for open-weight Chinese reasoning models following DeepSeek-R1’s January 2025 debut, and reinforced the argument, already gaining currency after R1, that competitive reasoning performance no longer required parameter counts or training budgets on the scale of the largest closed frontier models. A smaller, openly licensed model claiming near-parity with a much larger open one lowered the compute and cost bar for running a capable reasoning model locally, adding to pressure on both closed-model providers and larger open releases to justify their scale.

The gap between headline benchmark claims and harder evaluations showed up quickly: an ETH Zurich study published weeks later, grading full written proofs on the 2025 USA Mathematical Olympiad rather than final answers, found QwQ scored under 3% under expert human grading — in line with most other tested reasoning models, and far below the picture created by benchmarks measuring only correct final answers.