Download
4 variants available
GGUF (Quantized)
7640 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
(21)
Aug 4, 2026
MiniMax H3

2040 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
6340 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
20K0 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9K

MiniMax H3 is licensed by MiniMax under the MiniMax H3 Community License Agreement. That agreement’s Applicable Territory excludes the European Union, the United Kingdom, the Republic of Korea and the United States of America. Your use of H3 and of any H3 derivative is subject to that agreement and its Acceptable Use Policy.
MiniMax H3
**CIVITAI IS RENAMING EVERY DOWNLOADED FILE THE SAME THING. MAKE SURE YOU SAVE-AS AND NAME THE FILE SO YOU KNOW WHAT IT IS.**
Support
Everything here is free and stays free — the format spec, the nodes, the workflows, the cartridges, the LoRAs. If it saved you a night of debugging (it contains several hundred of mine), tips keep the 5090 warm:
🔁 Liberapay (recurring)
⚡ Or right here: the Civitai tip button on this page sends Buzz directly.
MiniMax-H3's 33B video+audio DiTs, quantized for 16-32 GB cards. Both flavors — fl2va (text / first-last-frame to video+audio) and ref2va (identity from reference images, videos, and voice clips) — quantized from the Comfy-Org bf16 releases with a first-class minimax_h3 architecture in the metadata.
The files
fl2va Q5_1 — 25.9 GB · 6.0 bits/weight · the 24-32 GB tier, near-Q6 quality. The demo on the companion workflow page rendered on this file.
fl2va Q4_0 — 19.9 GB · 4.5 bits/weight · the 16 GB tier (streams the overflow).
ref2va Q5_1 / Q4_0 — same sizes, for reference-identity workflows.
Coming after validation: Q3mix (17.3 GB) — real 3-bit K-quants on the tensors whose geometry allows it, classic types elsewhere. Named for what it is.
Quick fixes — read this first
“unknown model architecture: minimax_h3”
Why: ComfyUI-GGUF doesn't know this architecture out of the box.
Fix: Install ComfyUI-GGUF, then run the one-lineapply_gguf_arch_patch.pyfrom the attached loader pack (idempotent, takes two seconds), restart ComfyUI.Where is the text encoder? Where are the VAEs?
Why: A GGUF here contains the transformer and nothing else.
Fix: The text encoder is now quantized too — MiniMax-H3 Text Encoder GGUF (also on Hugging Face) (take themmprojfile with it, it is required for reference images AND for multi-shot chaining). VAEs come from Comfy-Org/MiniMax-H3. Load the GGUF withUnetLoaderGGUF— or with the one-dropdown dual-format loader in the companion workflow pack.It's 19.9 GB and I have a 16 GB card
Why: That is the wrong mental model — the file does not need to fit.
Fix: Nothing. ComfyUI streams what doesn't fit; you pay some speed, not failure. Measured reference: Q5_1 on an RTX 5090 runs ~22 GB resident, about 10 minutes for a 243-frame 544x960 shot at 20 steps.Which flavor do I want?
Why: Two different conditioning systems, two checkpoints.
Fix:fl2vafor text-to-video, image-to-video, and last-frame chaining (the multishot workflow).ref2vawhen identity comes from reference images or voice clips instead.
Why there is no Q6_K
K-quants (Q6_K, Q5_K, Q4_K, Q3_K…) require weight rows divisible by 256, and H3's hidden width is 2688 — which is not. Requesting one just quantizes something else with the wrong name on it. The ladder for pure quants here is the classic family: Q5_1 and Q4_0 above. One useful exception exists — the FFN down-projections ARE 256-divisible, which is what makes the mixed Q3mix tier possible at 17.3 GB.
Quantization notes
About 40% of this model is per-block adaln modulation; it is quantized here, exactly as the official int8 release quantizes it.
Kept at full precision: patch projections, time embedder, condition projection, final layer, token refiner.
Built with a patched llama.cpp adding
minimax_h3as a first-class architecture — the metadata says what the file actually is.