Alibaba releases Qwen2.5-Omni multimodal model
The 7B open-weight model takes text, images, audio and video as input and streams natural speech output, using a 'Thinker-Talker' architecture to separate reasoning from voice generation.
Open weights & ecosystem · Models & capabilities