release aidev NEEDS REVIEW SIG 1/5

MiniMax-H3 multimodal audio-video generation model published on Hugging Face

MiniMax published the MiniMax-H3 model repository on Hugging Face under the MiniMaxAI organisation. Pipeline tags indicate it is a diffusers-based image-text-to-video model supporting text-to-video, image-to-video, video-to-video, and synchronised audio-video generation across multiple modality combinations. No capability claims, benchmark results, or licensing details are asserted beyond the repository listing.

PUBLISHED2026-07-28
OBSERVED2026-08-11
AGE28d
SOURCES1
  • Published on Hugging Face as MiniMaxAI/MiniMax-H3
  • Pipeline: image-text-to-video (library: diffusers)
  • Tags indicate support for: text-to-video, image-to-video, image-text-to-video, video-to-video, text-to-audio-video, image-to-audio-video, image-text-to-audio-video, video-to-audio-video, audio-to-audio-video, audio-video-generation, multimodal, synchronized-audio-video, reference-to-audio-video
  • Diffusers pipeline: MiniMaxH3ModularPipeline
  • Region tag: us

COMMUNITY

No curated reactions recorded for this event. Facts and takes are kept in separate layers — community context is added by hand, never blended into the record above.