release aidev NEEDS REVIEW SIG 1/5

Xiaomi MiMo V2.5 multimodal model published

Xiaomi published MiMo V2.5, a multimodal model supporting vision, language, audio, video understanding, and agent capabilities with long context. Licensed under MIT. Tags indicate it handles text-generation, conversational interaction, and multimodal inputs including vision and audio.

PUBLISHED2026-04-27
OBSERVED2026-08-11
AGE4mo
SOURCES1
  • Pipeline: text-generation via Transformers
  • Tags: multimodal (vision-language, audio, video-understanding, agent, long-context)
  • Languages: en, zh
  • License: MIT
  • Format: safetensors, fp8
  • Hosted on Hugging Face under XiaomiMiMo organisation

COMMUNITY

No curated reactions recorded for this event. Facts and takes are kept in separate layers — community context is added by hand, never blended into the record above.