Qwen2-VL: vision-language model with 20+ minute video understanding
Qwen released Qwen2-VL, the latest vision-language model based on Qwen2. It achieves state-of-the-art on visual understanding benchmarks (MathVista, DocVQA, RealWorldQA, MTVQA) and can understand videos over 20 minutes for QA, dialog, and content creation.
PUBLISHED2024-08-28
OBSERVED2026-08-11
AGE1y
SOURCES1
- State-of-the-art on MathVista, DocVQA, RealWorldQA, MTVQA benchmarks
- Handles images of various resolutions and aspect ratios
- Can understand videos over 20 minutes for question answering, dialog, and content creation
COMMUNITY
No curated reactions recorded for this event. Facts and takes are kept in separate layers — community context is added by hand, never blended into the record above.