MiniMax H3 is an open-weight, omni-modal video model released by MiniMax on July 31, 2026. It generates video with native audio at up to 2K resolution and 15 seconds.
Key facts from MiniMax and the public model pages on fal.ai.
Understands text, images, video and audio together as context for a generation.
Generates sound in the same pass as the picture.
On fal: 480P and 768P native, 2K and 4K upscaled from a 768P base.
Weights are published on GitHub and model hubs, so you can run H3 on your own hardware.
H3 Max is a post-trained version of H3 released by fal.ai in August 2026.
Multimodal references, seven aspect ratios on fal, and local deployment.
Tuned for prompt adherence, visual quality and fast generation.
H3 on fal offers 2K and 4K upscaling; H3 Max goes up to 1080P.
Running H3 locally needs a large GPU setup. This site runs H3 Max in the cloud for you.