Download
1 variant available
This checkpoint includes a config file, download and place it along side the checkpoint.
5730 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
(24)
Aug 4, 2026
MiniMax H3

3660 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
2.5K0 1 2 3 4 5 6 7 8 9.0 1 2 3 4 5 6 7 8 9K
34.2K0 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9.0 1 2 3 4 5 6 7 8 9K
I have been experimenting with MiniMax H3 FL2VA as a pseudo-image generator in ComfyUI.
This is not a native text-to-image mode. H3 generates a short sequence, the video VAE decodes it into an image batch, and Image From Batch extracts one frame as the final still.
Portraits and cinematic images worked well, but the biggest surprise was typography and graphic design. I tested magazine layouts, posters, dashboards, infographics, charts, fantasy key art, and phone-style photography.
The outputs are not perfect, but H3 follows detailed art direction surprisingly well. It understands typography hierarchy, layout structure, palettes, charts, icons, ornament, and the relationship between text and imagery. The results become much stronger when the prompt defines the complete design instead of asking for a generic poster or infographic.
There can still be spelling mistakes, fake microtext, inaccurate chart data, and video-VAE artifacts, so the results need inspection. Still, this looks very useful for posters, covers, key art, presentation visuals, design exploration, and infographic drafts.
Parameter note
In my setup, these values produced the best results:
INT Length: 8
Image From Batch Index: 8Length controls the short sequence generated by H3. After decoding, Batch Index selects which frame is saved as the image.
The 8/8 combination is based only on my experiments. Preview the full decoded batch and test nearby values, since the cleanest frame may vary depending on the prompt, resolution, checkpoint, and node implementation.

