Qwen releases Qwen2.5-Omni end-to-end multimodal model
Qwen released Qwen2.5-Omni, a flagship multimodal model processing text, images, audio, and video with real-time streaming output via text and natural speech synthesis. The 7B model is openly available on Hugging Face, ModelScope, DashScope, and GitHub.
PUBLISHED2025-03-26
OBSERVED2026-08-11
AGE1y
SOURCES1
- End-to-end multimodal model accepting text, images, audio, and video inputs
- Generates real-time streaming responses in both text and synthesized speech
- 7B parameter variant openly released under standard channels
- Technical documentation published in accompanying paper
COMMUNITY
No curated reactions recorded for this event. Facts and takes are kept in separate layers — community context is added by hand, never blended into the record above.