MiMo-VL-7B-SFT-2508: SFT variant of Xiaomi's vision-language model
Xiaomi published the SFT-trained variant of its 7B vision-language model on Hugging Face. Shares the same Qwen2.5-VL architecture and MIT license as the RL variant, representing the supervised fine-tuning checkpoint of the same model family.
PUBLISHED2025-08-07
OBSERVED2026-08-11
AGE1y
SOURCES1
- Pipeline tag: image-text-to-text
- Library: transformers
- Architecture: qwen2_5_vl
- Parameters: 7B
- Tags: conversational, image-text-to-text
- Arxiv reference: 2506.03569
- License: MIT
- Text-generation-inference and endpoints compatible
COMMUNITY
No curated reactions recorded for this event. Facts and takes are kept in separate layers — community context is added by hand, never blended into the record above.