MiniMax H3 ComfyUI Guide: Setup, VRAM & GGUF (2026)

October 1, 2026

Yes, MiniMax H3 runs in ComfyUI natively. Support landed on August 3, 2026, the day MiniMax published the open weights. Update ComfyUI to 0.30.0 or later, open a MiniMax H3 template from the template library and download the files it asks for. The default pruned INT8 model is about 21 GB. Comfy says that with dynamic VRAM offloading, H3 can run on a GPU like the RTX 3060. Community GGUF quants go lower, down to about 9 GB.

Before you download anything, read the license section below — local use is not permitted everywhere.

Check the license first

MiniMax H3 is released under the MiniMax H3 Community License Agreement. Three points matter for local use:

  • Excluded territories. The license grants rights only outside the European Union, the United Kingdom, the Republic of Korea and the United States. Using the weights or their outputs in those territories is not authorized by the license; MiniMax asks people there to contact it about deployment.
  • Large companies. Commercial products or services with more than 20 million US dollars in revenue need separate written authorization from MiniMax.
  • Commercial use of local outputs. Comfy's documentation states that commercial use of locally generated outputs requires a MiniMax commercial license, available through Comfy. Generations on Comfy Cloud include commercial rights.

Read the full license before you rely on any of this. If you cannot or do not want to run H3 locally, you can use H3 Max online in your browser instead — no GPU or setup needed. If you need an open model you can run in the EU or US, see MiniMax H3 vs LTX-2.5.

What you need

  • ComfyUI 0.30.0 or later for text to video, image to video and reference to video. Some extras need newer versions: multiframe reference 0.34.0, Fun ControlNet Union and sparse attention 0.35.0, FastH3 0.36.0.
  • Disk space for a diffusion model, a text encoder and two VAEs (one for video, one for audio). About 35 GB for the smallest official set, about 52 GB for the default one.
  • System RAM to hold what does not fit in VRAM. Comfy's optimized setup has a memory footprint of about 42.5 GB, down from 123.6 GB in full precision.

Which model files to download

All official ComfyUI files are in the Comfy-Org/MiniMax-H3 repository. Pick FL2VA for text to video and image to video (first frame, last frame or both), and Ref2VA for reference to video.

PartFileSizeFolder
Diffusion model (default)minimax_h3_fl2va_pruned_int8_convrot.safetensors21.0 GBmodels/diffusion_models/
Diffusion model (smaller)minimax_h3_fl2va_pruned_w6a8.safetensors16.0 GBmodels/diffusion_models/
Diffusion model (higher VRAM)minimax_h3_fl2va_pruned_bf16.safetensors40.2 GBmodels/diffusion_models/
Text encoderqwen3vl_32b_minimax_h3_int8_convrot.safetensors27.1 GBmodels/text_encoders/
Text encoder (smaller)qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors15.7 GBmodels/text_encoders/
Video VAEminimax_h3_video_vae_int8_convrot.safetensors2.8 GBmodels/vae/
Audio VAEminimax_h3_audio_vae_fp32.safetensors0.6 GBmodels/vae/

For reference to video, swap the diffusion model for the matching minimax_h3_ref2va_… file. Comfy recommends the pruned checkpoints: the pruning replaces modulation weights with a lookup table, which shrinks memory without changing the output.

Set up the workflow step by step

  1. Update ComfyUI to 0.30.0 or later (or use Comfy Cloud).
  2. Open Template Library → Video and pick a MiniMax H3 template: video_minimax_h3_t2v (text to video), video_minimax_h3_i2v (image to video) or video_minimax_h3_r2v (reference to video).
  3. Download the models from the pop-up, or place the files from the table above into their folders.
  4. Set the resolution. The templates start at a fast preview size. For full quality at 16:9, set the Resolution Selector to 0.98 megapixels with a multiple of 32, which gives 1344×768 — H3's native 768 px short edge. Avoid 1.0 megapixels: it produces 1376×768, which exceeds the model's pixel cap.
  5. Keep the default sampling. The templates use the res_multistep sampler, the simple scheduler and 20 steps.
  6. Write your prompt and queue it. Clips run 4–15 seconds at 24 fps with 32 kHz stereo audio. Our 30 H3 Max prompts work as starting points.

Local output tops out at 768p. The 2K upscaling module (H3-Regenerate-2K) and the prompt-refinement layer (H3-Context-IR) that MiniMax uses in its hosted service are not part of the open release.

Steps and quality

  • Simple shots hold up at 12–16 steps.
  • High-frequency detail keeps improving up to about 50 steps.
  • Audio keeps improving after the picture has stopped changing, so do not cut steps too far if sound matters.
  • Reference drift at 20 steps: try 25 steps; short schedules weaken the reference.

Make it faster

  • Lightning LoRA. Turn on Enable Lightning LoRA (or the turbo_mode widget) in the template. Text to video and image to video then run in 8 steps; reference to video in 4 steps. The LoRA files are in the same Comfy-Org repository under loras/. To train your own, see our MiniMax H3 LoRA guide.
  • Sage Attention. Install a matching SageAttention wheel and the KJNodes pack, then place Patch Sage Attention KJ between UNETLoader and BasicGuider with sage_attention set to auto. Comfy reports roughly double the speed with minimal quality loss.
  • FastH3. A 4-step distilled version, supported from ComfyUI 0.36.0.

Low VRAM: community GGUF quants

GGUF builds of H3 come from the community, not from MiniMax or Comfy. They load through ComfyUI-GGUF with the UnetLoaderGGUF node.

RepositoryFilesNotes
joeygambino/MiniMax-H3-GGUFQ5_1 25.9 GB, Q4_0 19.9 GBQ5_1 for 24–32 GB cards, Q4_0 for 16 GB cards (streams the overflow). Pruned set ~40% smaller in MiniMax-H3-curve-GGUF
Abiray/MiniMax-H3-Pruned-GGUF8.9–21.6 GBPruned quants labeled Q3_K_M to Q8_0; Q4_K_M suggested for 16 GB

Two things to know:

  • "Unexpected architecture type in GGUF file: 'minimax_h3'". The file is fine; ComfyUI-GGUF does not list the architecture yet. Installing ComfyUI-H3-Multishot v1.5.2 or newer registers it automatically at startup.
  • K-quant labels. The joeygambino maintainer notes that true K-quants cannot be built for H3's original weights, because its hidden width (2688) is not divisible by 256. Treat "K" labels in other repositories with care and compare outputs yourself.

Streaming does not shrink the model; it moves weights between RAM and VRAM. On smaller cards a lower quant mainly costs quality, while streaming mainly costs speed.

Troubleshooting

ProblemFix
Out of memoryUse the pruned INT8 or W6A8 diffusion model and the NVFP4 text encoder, or a GGUF quant
GGUF architecture errorInstall ComfyUI-H3-Multishot v1.5.2+ (see above)
Morphing near the end of the clip or garbled text with Sage Attention on INT8Switch to the bf16 checkpoint and add a Model Attention Backend node set to comfy kitchen attention
Output looks softCheck the resolution is 0.98 MP (1344×768), not the template's preview size

MiniMax H3 locally vs H3 Max online

H3 Max cannot be run in ComfyUI. It is a post-trained version of H3 that is only available as a hosted model, so the open weights you download are the base H3. See H3 Max vs H3 for the differences, and MiniMax H3 for an overview of the base model.

MiniMax H3 in ComfyUIH3 Max on h3maxvideo.net
HardwareYour GPU, plus lots of RAM and diskNone — runs in the browser
SetupComfyUI 0.30.0+, ~35–52 GB of filesSign in with Google
ResolutionUp to 768p locally480P, 768P, 1080P
InputsText, first/last frame, multi-referenceText, first frame + optional last frame
LicenseCommunity license, territory limits applyOur Terms of Service
CostYour hardware and electricityCredits from $2.99 — see pricing

If you just want to see what H3 can do, start with the free credits on the H3 Max video generator, or pick H3 Max Turbo for half-price drafts.

FAQ

Is MiniMax H3 free to run in ComfyUI? The weights are free to download under the community license, within the licensed territories. Commercial use of local outputs needs a commercial license, according to Comfy.

What is the minimum GPU for MiniMax H3? Comfy says H3 runs on a GPU like the RTX 3060 with dynamic VRAM offloading, as long as you have enough system RAM. Larger GPUs mainly make it faster.

Where do I put the GGUF file? In ComfyUI/models/unet/, loaded with the UnetLoaderGGUF node from ComfyUI-GGUF. Official safetensors files go in models/diffusion_models/.

FL2VA or Ref2VA? FL2VA for text to video and first/last frame image to video. Ref2VA for generating from reference images, videos and audio.

Can I run H3 Max or H3 Max Turbo locally? No. Both are hosted models. Only the base MiniMax H3 has open weights.

Sources

Last updated: October 1, 2026. h3maxvideo.net is independent and not affiliated with MiniMax, Comfy or fal.ai.

H3Max Video

H3Max Video

MiniMax H3 ComfyUI Guide: Setup, VRAM & GGUF (2026) | H3Max Video