BAAI announces Wu Dao 2.0 at 1.75 trillion parameters
The Beijing Academy of Artificial Intelligence put the parameter count at ten times GPT-3's, a claim reported widely but never independently benchmarked.
- Models & capabilities
- Notable
The Beijing Academy of Artificial Intelligence (BAAI) unveiled Wu Dao 2.0, describing it as trained with 1.75 trillion parameters — roughly ten times the 175 billion of OpenAI’s GPT-3. BAAI, a state-backed research consortium in Beijing, presented the model as a multimodal system capable of generating text and images, holding conversations, understanding pictures and even composing poetry and recipes, trained on a reported 1.2 terabytes of English and Chinese text.
Wu Dao 2.0 was assembled from four component models rather than a single dense network: a Chinese text model BAAI called Wen Yuan, a multimodal model called Wen Lan, a generative language model called Wen Hui, and a mixture-of-experts model called Wen Su, each credited with its own efficiency or accuracy claims relative to prior systems such as BERT and RoBERTa. The mixture-of-experts design meant the 1.75 trillion figure counted parameters spread across many specialised sub-networks rather than all active on every input — the same architectural choice Google had used months earlier in its own trillion-parameter Switch Transformer.
No independent evaluation accompanied the announcement. The reported benchmarks came from BAAI itself, and no standard leaderboard score, code release or reproducible demonstration let outside researchers test the claims. Parameter count, in the absence of comparable benchmarking, functioned mainly as a headline number.
The announcement mattered less for what could be verified than for what it signalled: that the scaling race framed by GPT-3 was no longer an exclusively American or even exclusively industry pursuit, and that a Chinese state-funded lab could credibly claim to have built the largest model yet announced, parameter-count claims notwithstanding.