release aidev NEEDS REVIEW SIG 1/5

MiMo-VL-7B-SFT-2508: SFT variant of Xiaomi's vision-language model

Xiaomi published the SFT-trained variant of its 7B vision-language model on Hugging Face. Shares the same Qwen2.5-VL architecture and MIT license as the RL variant, representing the supervised fine-tuning checkpoint of the same model family.

PUBLISHED2025-08-07
OBSERVED2026-08-11
AGE1y
SOURCES1
  • Pipeline tag: image-text-to-text
  • Library: transformers
  • Architecture: qwen2_5_vl
  • Parameters: 7B
  • Tags: conversational, image-text-to-text
  • Arxiv reference: 2506.03569
  • License: MIT
  • Text-generation-inference and endpoints compatible

COMMUNITY

No curated reactions recorded for this event. Facts and takes are kept in separate layers — community context is added by hand, never blended into the record above.