Download
1 variant available
930 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
(6)
Aug 5, 2026
MiniMax H3
Critical fix. If you used the long-form MEMORY workflow in v1.1 or v1.2, it could not have worked. Update the node pack and restart ComfyUI fully — a browser refresh is not enough.
The bug
H3MultishotMemorySampler called vae_decode_audio() without importing it, raising a NameError at audio decode — after all sampling had finished. So it destroyed a completed render rather than failing fast, which is the worst possible moment to fall over.
Checking every commit that has touched that file, the import appears exactly once in all of them, in the other sampler. This node had never worked in any released version. It shipped broken in v1.1 and again in v1.2.
Reported with a correct diagnosis and the correct patch by Nebuluss in issue #1. Fixed and verified with a real two-shot render — 247 frames, audio stream present, and the embedded graph confirms the memory sampler itself produced it.
Also in this version: H3 Free Text Encoder
Unloads the text encoder before the DiT loads. This matters most for reference-to-video, where the Qwen3-VL encoder (~16.5 GB) and the ref2va DiT (~25.9 GB) together oversubscribe a 32 GB card. Measured on an RTX 5090, ref2va with one reference image at 124 frames:
encoder resident, DiT fully loaded: 90+ min, killed, no output
encoder evicted + --reserve-vram 7: 12.3 min, successBoth are needed. Eviction alone still let the DiT load completely at 25.2 GB with nothing left for activations.
The diagnostic that found this was power draw, not utilisation. The GPU reported 98% "utilisation" at 146 W — a 5090 doing real diffusion pulls 400 W+. 98% at 146 W is a card spinning on memory transfers. With headroom it runs at 410 W. Full residency is not the goal; total residency under the card size is.
What changed in how this gets released
An import that survived two releases means the process failed, not just the code. Two checks now gate every package:
Provenance, not existence. My evidence file claimed the MEMORY workflow had a render behind it — pointing at a file actually produced by the other sampler. The gate confirmed the file existed and never asked what made it. It now requires the workflow's own node class to appear in the render's embedded ComfyUI graph.
A static import check over every shipped
.py, so a call that parses fine butNameErrors at runtime cannot ship again.
Everything from v1.2 is still here
Keyframes at any position — up to 6 anchors by fraction or frame index, not just first/last. Measured: an anchor requested at frame 121 landed at frame 122 of 243, reached by continuous motion with no cut.
H3 Condition Strength and H3 Reference Audio (stereo guard — a mono reference clip crashes the sampler with an unhelpful shape mismatch).
Image-to-video, the ~4x encoder-eviction speedup for multishot, crossfaded audio seams, and the self-repairing script parser.
Links
Support
Everything I publish is free and stays free. If it saved you a night of debugging, tips keep the 5090 warm: Ko-fi · GitHub Sponsors · Liberapay.
Show more

690 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
1760 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9

Support
Everything here is free and stays free — the format spec, the nodes, the workflows, the cartridges, the LoRAs. If it saved you a night of debugging (it contains several hundred of mine), tips keep the 5090 warm:
🔁 Liberapay (recurring)
⚡ Or right here: the Civitai tip button on this page sends Buzz directly.
Three ways to drive MiniMax-H3, and a VRAM fix that makes it usable on a 32 GB card. Chain shots from a script into one long piece; pin keyframes anywhere in a clip; or drive identity from reference images, video and voice. All in one node pack, all producing video and audio.
Keyframes at any position
Stock ComfyUI pins H3 keyframes to the first and last frame only and raises only first/last keyframe anchors are supported for anything else.
That is a positional-maths limit, not a model limit. Both stock cases are the same expression, because sum(_video_t_spans(latent_t)) == FRAME_RESCALE * frame_count:
cond_t = text_len + FRAME_RESCALE * pixel_index— which is defined for every frame, not just the two endpoints. So you can hand H3 up to six anchor images and say when each one should happen, as fractions (0, 0.5, 1) or absolute frame indices.
Measured on an RTX 5090, 243 frames, one anchor at pixel frame 121: the rendered frame most resembling the anchor image was frame 122 — the requested position, off by one — arrived at by continuous motion with no cut (peak frame-to-frame change 2.3× the median), and the audio ran unbroken straight through it. A three-anchor run at 0 / 0.5 / 1 landed the second and third on frames 121 and 242 exactly.
The patch is applied in memory — it does not edit any ComfyUI file. It self-tests against the stock formula before committing and rolls itself back if first/last positions do not reproduce exactly, so a future ComfyUI change degrades to “interior anchors unavailable” rather than to broken renders.
It moves, or it cuts — and your images decide which
Anchor images with a plausible camera path between them (same place, different angle or framing) make H3 interpolate: a real move that arrives on time. Images with no possible path — a kitchen and a diner — make it cut, then hold.
That is the model being sensible, not a limitation of the node: stock first/last does exactly the same thing when given such a pair. And the cut case earns its keep, because it is a timed shot change inside a single generation — one generation means one continuous audio stream, so the voice does not get re-derived and there is no seam to hide.
Which workflow
You wantUseA script of several shots, chained into one pieceH3_Multishot_AIO.jsonSpecific frames at specific times, one continuous takeH3_Keyframes.json2–5 minutes without identity driftingH3_Multishot_MEMORY.json
A reference-to-video workflow is in progress and not in this release yet. Worth knowing for when it lands: references and keyframes are mutually exclusive. This is ComfyUI core behaviour, not a choice made here: model_base.py assigns cond_video_latents for references, discarding keyframe latents while the keyframe layout rows survive — the packed sequence then desyncs into a shape-mismatch crash. There is no “reference images plus start frame” mode. Pick one per shot.
~4× faster on 32 GB cards
The text encoder is evicted before sampling. The Qwen3-VL encoder (~16.5 GB even at Q4) and the H3 DiT (~25 GB) do not co-fit on a 32 GB card, so the DiT was loading partially and streaming ~19 GB from system RAM on every sampling step. If you have ever seen this in your log:
loaded partially; 6423 MB usable, 5847 MB loaded, 19363 MB offloadedthat was it. Measured on an RTX 5090: ~60 min → ~15 min, same render. Every sampler in this pack does it, including the keyframes node, and prints TE evicted; NN.N GB free for the DiT.
Quick fixes — read this first
Red / missing nodes when a workflow loads
Why: pack not installed, or ComfyUI not restarted.
Fix: install ComfyUI-H3-Multishot (Manager > Install via Git URL), restart, then hard-refresh the browser tab — the frontend caches node definitions.GGUF errors with “unknown model architecture”
Why: ComfyUI-GGUF does not know MiniMax-H3 out of the box.
Fix: runpython apply_gguf_arch_patch.pyfrom the pack folder (one line, idempotent), restart. This is for the DiT only — the text encoder is Qwen3-VL and needs no patch.Reference audio crashes the sampler with a shape mismatch
Why: your clip is mono. The audio VAE encodes[B, 2, L]and the layout reserves exactly two channels, so a mono reference produces half the rows it reserved and dies deep inside the model with no useful message. Nothing in stock converts it.
Fix: the H3 Reference Audio node — forces stereo 32 kHz and trims length. It is already wired in the hard-mode workflow.The model ignores my reference image
Why: reference blocks are labelled in the prompt and the numbering is 1-based while the input slots are 0-based —ref_image_0is<Picture 1>. If you never name it in the text, the model has no reason to bind it.
Fix: write<Picture 1> is the woman. <Picture 2> is the room.Note a reference video with a soundtrack consumes an<Audio j>ordinal before your standalone clips.Multi-shot ignores my reference image entirely
Why: missingmmproj. It is required for chaining, not just for reference images — chaining feeds the previous shot's last frame through the encoder's vision path.
Fix: download the-mmprojfile alongside the encoder and keep both filenames exactly as downloaded, in the same folder. The loader pairs them by name.Length change errors out
Why: H3 hard constraint — frame counts live on a 17k+5 grid.
Fix: 226, 243, 260… the widget steps by 17 so it keeps you legal. 243 ≈ 10s; 362 ≈ 15s is the trained ceiling.Speech turns to gibberish
Why: usually an under-filled shot, not an over-long one. Speech runs about 2.5 words/second, so a 243-frame shot wants roughly 22–25 spoken words; give it eight and the model invents sound to fill the dead air.
Fix: match dialogue length to shot length, and if you want silence, script it (“she listens, saying nothing”).An object morphs into something else at a seam
Why: chaining hands each shot the previous final frame. If shot 1 ends on the cameraman holding his camcorder and shot 2 is filmed FROM that camcorder, the model must explain a device in a hand that should not be in frame — so it invents one.
Fix: end every shot on what the NEXT shot expects to see.Out of VRAM, or renders crawl
Why: 33B of weights — and see the eviction section above.
Fix: use the GGUFs, Q5_1 for 24–32 GB, Q4_0 for 16 GB. The file does not need to fit in VRAM; ComfyUI streams the overflow. Expect ~10 min per 10s shot on a 5090-class card.
Writing a script — this is most of the quality
The identity lock is description density, not assertion. Every shot is an independent conditioning pass: the model rebuilds the person from your text each time. Writing “the same woman, same face, same wardrobe” asserts continuity without supplying what is needed to rebuild it, and the face drifts. Re-describing 6–8 concrete attributes verbatim in every shot is what actually holds it:
She is an attractive American woman in her mid twenties with warm hazel
eyes, a friendly confident smile, light freckles, shoulder-length auburn
hair tucked behind one ear, small gold stud earrings, and a relaxed
sage-green blouse. Her voice is a clear warm young woman's voice in a
casual American accent.Do the same for the voice: one short concrete line, repeated verbatim. Flowing prose beats SHOT: / Audio: labels.
What you need
ComfyUI v0.30.0+ (native MiniMax H3 support)
The node pack: ComfyUI-H3-Multishot
DiT weights: MiniMax-H3 GGUF (or on Hugging Face) (Q5_1 / Q4_0, both flavours) — or the originals
Text encoder: MiniMax-H3 Text Encoder GGUF (or on Hugging Face) — take the mmproj file too
VAEs: Comfy-Org/MiniMax-H3
For GGUF: ComfyUI-GGUF + the included one-line patch
A word on expectations
MiniMax-H3 is a 33B joint audio+video model and this pack started days after the weights landed. It works, and the measurements on this page are from real renders on one consumer GPU — but you may still need to tune to YOUR machine. If you get stuck, comment here or open a GitHub issue. I answer.
Everything else I've published
Support
Everything I publish is free and stays free. If it saved you a night of debugging, tips keep the 5090 warm: Ko-fi · GitHub Sponsors · Liberapay.