# MiMo-VL-7B-SFT-2508: SFT variant of Xiaomi's vision-language model

> Xiaomi published the SFT-trained variant of its 7B vision-language model on Hugging Face. Shares the same Qwen2.5-VL architecture and MIT license as the RL variant, representing the supervised fine-tuning checkpoint of the same model family.

| | |
|---|---|
| **Tool** | Xiaomi MiMo |
| **Version** | — |
| **Kind** | release |
| **Published** | 2025-08-07 |
| **Observed** | 2026-08-11 |
| **Significance** | 1/5 |
| **Breaking** | no |
| **Categories** | feature, model-support |

> **Note:** this classification is below our confidence threshold and is pending human review.

## What changed


- Pipeline tag: image-text-to-text
- Library: transformers
- Architecture: qwen2_5_vl
- Parameters: 7B
- Tags: conversational, image-text-to-text
- Arxiv reference: 2506.03569
- License: MIT
- Text-generation-inference and endpoints compatible


## Sources

- [hf_model](https://huggingface.co/XiaomiMiMo/MiMo-VL-7B-SFT-2508) — retrieved 2026-08-11


## Community

_No curated reactions recorded. Facts and community takes are kept in separate layers
and never blended._

---
Canonical: https://changelogs.info/xiaomi-mimo/mimo-vl-7b-sft-2508-sft-variant-of-xiaomi-s-vision-language-model
Entity: https://changelogs.info/xiaomi-mimo
Event ID: `evt_2025-08-07_xiaomi-mimo_xiaomimimo-mimo-vl-7b-sft-2508`
Licence: event synthesis © changelogs.info, CC BY 4.0. Linked sources belong to their vendors.
