Download
1 variant available
This checkpoint includes a config file, download and place it along side the checkpoint.
3460 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
(19)
Aug 28, 2026
MiniMax H3
The same workflow with an 8 step turbo lora added, so a clip takes roughly half as long. The picture is very close to the standard build at the same seed. Start here if you are testing ideas.
Show more

1.1M0 1 2 3 4 5 6 7 8 9.0 1 2 3 4 5 6 7 8 9M
31.4K0 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9.0 1 2 3 4 5 6 7 8 9K
163.5K0 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9.0 1 2 3 4 5 6 7 8 9K

MiniMax H3 is licensed by MiniMax under the MiniMax H3 Community License Agreement. That agreementβs Applicable Territory excludes the European Union, the United Kingdom, the Republic of Korea and the United States of America. Your use of H3 and of any H3 derivative is subject to that agreement and its Acceptable Use Policy.
MiniMax H3
Hand it one photo of a person. It gives you back a completely different shot, in a place you never photographed, with that same person in it. Your photo never shows up in the result. It is read, and then set aside.
That is not the same thing as image to video, even though almost everyone starts by assuming it is. Image to video takes your picture and makes it move. This reads your picture, puts it down, and writes a new scene from scratch with that face in it.
π¬ Watch it built step by step: the full walkthrough on YouTube. Every setting, in order, on a normal home machine.
Two files in this post
Standard build - 20 steps. The one to start with if you want the safest result.
Turbo build - 8 steps, and one extra file to download. On my machine a 10 second clip took 153 seconds here against 349 seconds on the standard build. Same person, same picture quality. I ran both on the same seed before putting this up.
What you need
minimax_h3_ref2va_pruned_int8_convrot.safetensors β
ComfyUI/models/diffusion_models/qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors β
ComfyUI/models/text_encoders/minimax_h3_video_vae_fp16.safetensors β
ComfyUI/models/vae/minimax_h3_audio_vae_fp32.safetensors β
ComfyUI/models/vae/minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors β
ComfyUI/models/loras/(turbo build only)
The folder is the part people get wrong. A correct file in the wrong folder looks exactly like a missing model, and ComfyUI will not tell you which.
Two things worth knowing
The prompt does far more work here than in the other H3 workflows. You are describing a whole scene, not a movement, because nothing about the scene comes from your photo.
Both decoders are needed. One is for the picture and one is for the sound, and leaving either out fails in a way that does not obviously say so.
I wrote up every setting, what each one does and the mistakes that cost me time, in the full written guide. It is free and there is no card.
There is also a paid course that goes further with this model if you want it. Either way the workflows here are the whole thing, not a trimmed version.

