# Qwen releases Qwen2.5-Omni end-to-end multimodal model

> Qwen released Qwen2.5-Omni, a flagship multimodal model processing text, images, audio, and video with real-time streaming output via text and natural speech synthesis. The 7B model is openly available on Hugging Face, ModelScope, DashScope, and GitHub.

| | |
|---|---|
| **Tool** | Qwen |
| **Version** | — |
| **Kind** | announcement |
| **Published** | 2025-03-26 |
| **Observed** | 2026-08-11 |
| **Significance** | 4/5 |
| **Breaking** | no |
| **Categories** | new-model, capability, model-support |

## What changed


- End-to-end multimodal model accepting text, images, audio, and video inputs
- Generates real-time streaming responses in both text and synthesized speech
- 7B parameter variant openly released under standard channels
- Technical documentation published in accompanying paper


## Sources

- [blog_rss](https://qwenlm.github.io/blog/qwen2.5-omni/) — retrieved 2026-08-11


## Community

_No curated reactions recorded. Facts and community takes are kept in separate layers
and never blended._

---
Canonical: https://changelogs.info/qwen/qwen-releases-qwen2-5-omni-end-to-end-multimodal-model
Entity: https://changelogs.info/qwen
Event ID: `evt_2025-03-26_qwen_qwen2-5-omni-see-hear-talk-write-do-it-all`
Licence: event synthesis © changelogs.info, CC BY 4.0. Linked sources belong to their vendors.
