announcement aidev SIG 4/5

Qwen2-VL: vision-language model with 20+ minute video understanding

Qwen released Qwen2-VL, the latest vision-language model based on Qwen2. It achieves state-of-the-art on visual understanding benchmarks (MathVista, DocVQA, RealWorldQA, MTVQA) and can understand videos over 20 minutes for QA, dialog, and content creation.

PUBLISHED2024-08-28
OBSERVED2026-08-11
AGE1y
SOURCES1
  • State-of-the-art on MathVista, DocVQA, RealWorldQA, MTVQA benchmarks
  • Handles images of various resolutions and aspect ratios
  • Can understand videos over 20 minutes for question answering, dialog, and content creation

COMMUNITY

No curated reactions recorded for this event. Facts and takes are kept in separate layers — community context is added by hand, never blended into the record above.