# MiniMax-H3 multimodal audio-video generation model published on Hugging Face

> MiniMax published the MiniMax-H3 model repository on Hugging Face under the MiniMaxAI organisation. Pipeline tags indicate it is a diffusers-based image-text-to-video model supporting text-to-video, image-to-video, video-to-video, and synchronised audio-video generation across multiple modality combinations. No capability claims, benchmark results, or licensing details are asserted beyond the repository listing.

| | |
|---|---|
| **Tool** | MiniMax |
| **Version** | — |
| **Kind** | release |
| **Published** | 2026-07-28 |
| **Observed** | 2026-08-11 |
| **Significance** | 1/5 |
| **Breaking** | no |
| **Categories** | — |

> **Note:** this classification is below our confidence threshold and is pending human review.

## What changed


- Published on Hugging Face as `MiniMaxAI/MiniMax-H3`
- Pipeline: image-text-to-video (library: diffusers)
- Tags indicate support for: text-to-video, image-to-video, image-text-to-video, video-to-video, text-to-audio-video, image-to-audio-video, image-text-to-audio-video, video-to-audio-video, audio-to-audio-video, audio-video-generation, multimodal, synchronized-audio-video, reference-to-audio-video
- Diffusers pipeline: `MiniMaxH3ModularPipeline`
- Region tag: us


## Sources

- [hf_model](https://huggingface.co/MiniMaxAI/MiniMax-H3) — retrieved 2026-08-11


## Community

_No curated reactions recorded. Facts and community takes are kept in separate layers
and never blended._

---
Canonical: https://changelogs.info/minimax/minimax-h3-multimodal-audio-video-generation-model-published-on-hugging-face
Entity: https://changelogs.info/minimax
Event ID: `evt_2026-07-28_minimax_minimaxai-minimax-h3`
Licence: event synthesis © changelogs.info, CC BY 4.0. Linked sources belong to their vendors.
