Xiaomi MiMo V2.5 multimodal model published
Xiaomi published MiMo V2.5, a multimodal model supporting vision, language, audio, video understanding, and agent capabilities with long context. Licensed under MIT. Tags indicate it handles text-generation, conversational interaction, and multimodal inputs including vision and audio.
PUBLISHED2026-04-27
OBSERVED2026-08-11
AGE4mo
SOURCES1
- Pipeline: text-generation via Transformers
- Tags: multimodal (vision-language, audio, video-understanding, agent, long-context)
- Languages: en, zh
- License: MIT
- Format: safetensors, fp8
- Hosted on Hugging Face under XiaomiMiMo organisation
COMMUNITY
No curated reactions recorded for this event. Facts and takes are kept in separate layers — community context is added by hand, never blended into the record above.