MiniMax-H3 multimodal audio-video generation model published on Hugging Face
MiniMax published the MiniMax-H3 model repository on Hugging Face under the MiniMaxAI organisation. Pipeline tags indicate it is a diffusers-based image-text-to-video model supporting text-to-video, image-to-video, video-to-video, and synchronised audio-video generation across multiple modality combinations. No capability claims, benchmark results, or licensing details are asserted beyond the repository listing.
PUBLISHED2026-07-28
OBSERVED2026-08-11
AGE28d
SOURCES1
- Published on Hugging Face as
MiniMaxAI/MiniMax-H3 - Pipeline: image-text-to-video (library: diffusers)
- Tags indicate support for: text-to-video, image-to-video, image-text-to-video, video-to-video, text-to-audio-video, image-to-audio-video, image-text-to-audio-video, video-to-audio-video, audio-to-audio-video, audio-video-generation, multimodal, synchronized-audio-video, reference-to-audio-video
- Diffusers pipeline:
MiniMaxH3ModularPipeline - Region tag: us
COMMUNITY
No curated reactions recorded for this event. Facts and takes are kept in separate layers — community context is added by hand, never blended into the record above.