You can train LoRAs for MiniMax H3 on fal.ai or on your own GPU, and run them with H3's LoRA endpoints or in ComfyUI. fal has four H3 trainers — text to video, image to video, first/last frame and reference to video — that cost $5–15 per 1,000 training steps and return a .safetensors file. One thing to know up front: LoRAs work with the base MiniMax H3 model, not with H3 Max or H3 Max Turbo, which do not accept LoRA inputs.
What a MiniMax H3 LoRA does
A LoRA is a small add-on file that teaches H3 something specific — a character, a product, a visual style or a type of motion — without retraining the whole 33B model. Because H3 generates sound with the picture, an H3 LoRA can also learn audio: clips with a soundtrack teach the matching sound, and clips without audio train against silence.
The four fal trainers
| Trainer | Endpoint | Best for | Price |
|---|---|---|---|
| Text to video | minimax/h3/t2v/trainer | Styles and motion from prompts alone | $0.005 per step ($5 per 1,000) |
| Image to video | minimax/h3/i2v/trainer | Animating a starting image | $0.01 per step ($10 per 1,000) |
| First/last frame | minimax/h3/flf2v/trainer | Keyframe-driven shots | $0.01 per step ($10 per 1,000) |
| Reference to video | minimax/h3/ref2va/trainer | A consistent subject from reference images | $0.015 per step ($15 per 1,000) |
The default run is 2,000 steps, so a default text-to-video training costs $10 and a reference-to-video training $30.
Prepare the dataset
- Videos only. Upload a
.zipof.mp4,.mov,.avior.mkvclips. Image-only and mixed image/video archives are rejected. - At least 10 clips, more is better.
- Captions go in text files with the same base name as each clip (
clip01.mp4→clip01.txt). Describe what happens in the clip, including sound. - Frame rate defaults to 24 fps; clips longer than 30 seconds are split into scenes automatically.
- Check before you pay for a long run:
debug_datasetreturns your preprocessed data for inspection, andstrict_datasetrejects clips with missing captions or unreadable audio instead of silently falling back.
Training settings
| Setting | Default | Notes |
|---|---|---|
number_of_steps | 2000 | 1–15,000; cost scales with steps |
rank | 32 | 8, 16, 32, 64 or 128; higher means more capacity and memory |
learning_rate | 0.0002 | Higher learns faster but overfits sooner |
number_of_frames | 73 | 22–124, and frames % 17 must equal 5 (22, 39, 56, 73, 90, 107, 124) |
resolution | medium | low, medium or high |
aspect_ratio | 16:9 | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 |
trigger_phrase | empty | Prepended to every caption; use it in prompts later |
The image-to-video trainer conditions on the first frame half of the time by default (first_frame_conditioning_p 0.5); the reference trainer uses the subject's reference images 90% of the time.
Run your LoRA
fal's LoRA inference endpoints are minimax/h3/text-to-video/lora and minimax/h3/image-to-video/lora. Pass the trained file in loras with a scale:
{
"prompt": "my_style, a lighthouse on a cliff at dusk, slow aerial pull-back, waves and wind",
"loras": [
{ "path": "https://your-storage/lora.safetensors", "scale": 1 }
]
}These endpoints run the base H3 model at 480P ($0.0625/s), 768P ($0.075/s), 2K ($0.1625/s) or 4K ($0.20/s); 2K and 4K are upscaled from a 768P result. Clips are 5–15 seconds. To compare runs fairly, keep the prompt and seed the same and change only the LoRA or its scale.
In ComfyUI, load the .safetensors file with a LoRA loader node in a MiniMax H3 workflow — see our MiniMax H3 ComfyUI guide for the base setup.
Train locally
Community trainers have added H3 support. One Japanese creator reports training character LoRAs on a 12 GB RTX 4070 with musubi-tuner (experimental H3 support, including one-frame training from still images) and with ai-toolkit (needed code changes), using quantized models and CPU offloading. Treat these as community reports, not official requirements.
Local training uses the open weights, so the MiniMax H3 license applies — including its exclusion of the EU, UK, US and South Korea.
"Turbo LoRA" is not H3 Max Turbo
Two different things share the word "turbo":
- H3 Turbo / Lightning LoRAs are speed LoRAs for base H3 in ComfyUI. Comfy publishes
minimax_h3_fl2v_turbo_8stepand…_4stepfiles that cut generation from 20 steps to 8 or 4. - H3 Max Turbo is a separate hosted model on fal — a faster, half-price version of H3 Max. It is not a LoRA and does not take LoRAs. See H3 Max Turbo.
Do LoRAs work with H3 Max?
No. H3 Max and H3 Max Turbo are hosted models whose APIs have no LoRA input, and all fal trainers target base H3. If you want H3 Max quality without training anything, write a precise prompt — our H3 Max prompt guide covers the format — or use image to video to lock the look with a first frame.
You can try both modes on this site with free credits: H3 Max text to video and image to video.
FAQ
How much does it cost to train a MiniMax H3 LoRA on fal? $0.005 to $0.015 per step depending on the trainer. At the default 2,000 steps that is $10 (text to video) to $30 (reference to video).
How many clips do I need? At least 10. More varied clips generally give a more robust LoRA.
Can I train an H3 LoRA from images? Not on fal — its trainers accept videos only. Community tools such as musubi-tuner report experimental one-frame training from still images.
Does an H3 LoRA learn audio? Yes. Clips with a soundtrack teach the matching sound; clips without audio train against silence.
Can I use my LoRA on h3maxvideo.net? No. This site runs H3 Max and H3 Max Turbo, which do not support LoRAs.
Sources: fal.ai trainer pages for text to video, image to video, first/last frame and reference to video; fal: H3 text to video with LoRA; Comfy-Org/MiniMax-H3 files. Last updated: October 1, 2026.
h3maxvideo.net is independent and not affiliated with MiniMax or fal.ai.

