Timeline

Baidu releases ERNIE 5.0 preview

A natively omni-modal model jointly trained on text, images, audio and video, which Baidu presented as competitive with GPT-5 and Gemini 2.5 Pro on its own benchmark slides.

  • Models & capabilities
  • Colour

Baidu unveiled a preview of ERNIE 5.0 at its annual Baidu World conference and put it into public testing through its ERNIE Bot app and, for enterprise customers, its Qianfan cloud platform. The company described the model as “natively omni-modal,” trained jointly on text, images, audio and video rather than bolting vision or speech onto a text-first architecture.

Baidu’s own benchmark slides, shown at its Baidu World 2025 conference, placed ERNIE 5.0 ahead of or level with OpenAI’s GPT-5 and Google’s Gemini 2.5 Pro on multimodal reasoning, document understanding and image-based question answering — comparisons Baidu selected and published itself, not independently verified results. The model’s parameter count and training compute were not disclosed.

The release continued a pattern among Chinese labs of pairing frontier-scale multimodal claims with free or low-cost public access, positioning ERNIE 5.0 against the leading US models even as Baidu withheld its parameter count and training details.