# Qwen VLo: unified multimodal understanding and generation model

> Qwen introduced VLo, a unified multimodal model that both understands images and generates high-quality visual content based on that understanding. It builds on the Qwen2.5 VL line, extending from perception-only to perception-plus-creation.

| | |
|---|---|
| **Tool** | Qwen |
| **Version** | — |
| **Kind** | announcement |
| **Published** | 2025-06-26 |
| **Observed** | 2026-08-11 |
| **Significance** | 3/5 |
| **Breaking** | no |
| **Categories** | new-model, capability |

## What changed


- Unified multimodal understanding and generation in a single model
- Described as evolving from the Qwen2.5 VL line
- Capable of generating visual content based on understanding of image inputs
- No architecture details, benchmark results, or release date provided in the announcement


## Sources

- [blog_rss](https://qwenlm.github.io/blog/qwen-vlo/) — retrieved 2026-08-11


## Community

_No curated reactions recorded. Facts and community takes are kept in separate layers
and never blended._

---
Canonical: https://changelogs.info/qwen/qwen-vlo-unified-multimodal-understanding-and-generation-model
Entity: https://changelogs.info/qwen
Event ID: `evt_2025-06-26_qwen_qwen-vlo-from-understanding-the-world-to-depicting-it`
Licence: event synthesis © changelogs.info, CC BY 4.0. Linked sources belong to their vendors.
