# MiMo-VL-7B-RL-2508: RL-trained vision-language model from Xiaomi

> Xiaomi published the RL-trained variant of its 7B vision-language model on Hugging Face. Based on the Qwen2.5-VL architecture, the model supports image-text-to-text conversational tasks under an MIT license.

| | |
|---|---|
| **Tool** | Xiaomi MiMo |
| **Version** | — |
| **Kind** | release |
| **Published** | 2025-08-07 |
| **Observed** | 2026-08-11 |
| **Significance** | 1/5 |
| **Breaking** | no |
| **Categories** | feature |

> **Note:** this classification is below our confidence threshold and is pending human review.

## What changed


- Pipeline tag: image-text-to-text
- Library: transformers
- Architecture: qwen2_5_vl
- Parameters: 7B
- Tags: conversational, image-text-to-text
- Arxiv reference: 2506.03569
- License: MIT
- Text-generation-inference and endpoints compatible


## Sources

- [hf_model](https://huggingface.co/XiaomiMiMo/MiMo-VL-7B-RL-2508) — retrieved 2026-08-11


## Community

_No curated reactions recorded. Facts and community takes are kept in separate layers
and never blended._

---
Canonical: https://changelogs.info/xiaomi-mimo/mimo-vl-7b-rl-2508-rl-trained-vision-language-model-from-xiaomi
Entity: https://changelogs.info/xiaomi-mimo
Event ID: `evt_2025-08-07_xiaomi-mimo_xiaomimimo-mimo-vl-7b-rl-2508`
Licence: event synthesis © changelogs.info, CC BY 4.0. Linked sources belong to their vendors.
