# Qwen2-VL: vision-language model with 20+ minute video understanding

> Qwen released Qwen2-VL, the latest vision-language model based on Qwen2. It achieves state-of-the-art on visual understanding benchmarks (MathVista, DocVQA, RealWorldQA, MTVQA) and can understand videos over 20 minutes for QA, dialog, and content creation.

| | |
|---|---|
| **Tool** | Qwen |
| **Version** | — |
| **Kind** | announcement |
| **Published** | 2024-08-28 |
| **Observed** | 2026-08-11 |
| **Significance** | 4/5 |
| **Breaking** | no |
| **Categories** | new-model, feature, capability |

## What changed


- State-of-the-art on MathVista, DocVQA, RealWorldQA, MTVQA benchmarks
- Handles images of various resolutions and aspect ratios
- Can understand videos over 20 minutes for question answering, dialog, and content creation


## Sources

- [blog_rss](https://qwenlm.github.io/blog/qwen2-vl/) — retrieved 2026-08-11


## Community

_No curated reactions recorded. Facts and community takes are kept in separate layers
and never blended._

---
Canonical: https://changelogs.info/qwen/qwen2-vl-vision-language-model-with-20-minute-video-understanding
Entity: https://changelogs.info/qwen
Event ID: `evt_2024-08-28_qwen_qwen2-vl-to-see-the-world-more-clearly`
Licence: event synthesis © changelogs.info, CC BY 4.0. Linked sources belong to their vendors.
