release aidev NEEDS REVIEW SIG 1/5

MiMo-VL-7B-RL-2508: RL-trained vision-language model from Xiaomi

Xiaomi published the RL-trained variant of its 7B vision-language model on Hugging Face. Based on the Qwen2.5-VL architecture, the model supports image-text-to-text conversational tasks under an MIT license.

PUBLISHED2025-08-07
OBSERVED2026-08-11
AGE1y
SOURCES1
  • Pipeline tag: image-text-to-text
  • Library: transformers
  • Architecture: qwen2_5_vl
  • Parameters: 7B
  • Tags: conversational, image-text-to-text
  • Arxiv reference: 2506.03569
  • License: MIT
  • Text-generation-inference and endpoints compatible

COMMUNITY

No curated reactions recorded for this event. Facts and takes are kept in separate layers — community context is added by hand, never blended into the record above.