# Xiaomi MiMo V2.5 multimodal model published

> Xiaomi published MiMo V2.5, a multimodal model supporting vision, language, audio, video understanding, and agent capabilities with long context. Licensed under MIT. Tags indicate it handles text-generation, conversational interaction, and multimodal inputs including vision and audio.

| | |
|---|---|
| **Tool** | Xiaomi MiMo |
| **Version** | — |
| **Kind** | release |
| **Published** | 2026-04-27 |
| **Observed** | 2026-08-11 |
| **Significance** | 1/5 |
| **Breaking** | no |
| **Categories** | capability |

> **Note:** this classification is below our confidence threshold and is pending human review.

## What changed


- Pipeline: text-generation via Transformers
- Tags: multimodal (vision-language, audio, video-understanding, agent, long-context)
- Languages: en, zh
- License: MIT
- Format: safetensors, fp8
- Hosted on Hugging Face under XiaomiMiMo organisation


## Sources

- [hf_model](https://huggingface.co/XiaomiMiMo/MiMo-V2.5) — retrieved 2026-08-11


## Community

_No curated reactions recorded. Facts and community takes are kept in separate layers
and never blended._

---
Canonical: https://changelogs.info/xiaomi-mimo/xiaomi-mimo-v2-5-multimodal-model-published
Entity: https://changelogs.info/xiaomi-mimo
Event ID: `evt_2026-04-27_xiaomi-mimo_xiaomimimo-mimo-v2-5`
Licence: event synthesis © changelogs.info, CC BY 4.0. Linked sources belong to their vendors.
