MiMo-VL-7B-RL-2508: RL-trained vision-language model from Xiaomi
Xiaomi published the RL-trained variant of its 7B vision-language model on Hugging Face. Based on the Qwen2.5-VL architecture, the model supports image-text-to-text conversational tasks under an MIT license.
PUBLISHED2025-08-07
OBSERVED2026-08-11
AGE1y
SOURCES1
- Pipeline tag: image-text-to-text
- Library: transformers
- Architecture: qwen2_5_vl
- Parameters: 7B
- Tags: conversational, image-text-to-text
- Arxiv reference: 2506.03569
- License: MIT
- Text-generation-inference and endpoints compatible
COMMUNITY
No curated reactions recorded for this event. Facts and takes are kept in separate layers — community context is added by hand, never blended into the record above.