Updated: Aug 4, 2026
base modelWorkflows: https://huggingface.co/molbal/MiniMax-H3-GGUF/tree/main/workflows
Loader nodes: https://github.com/molbal/ComfyUI-GGUF / https://registry.comfy.org/publishers/molbal/nodes/comfyui-gguf-reboot
MiniMax H3
System Overview
MiniMax H3 is a general-purpose, omni-modal generative system. It supports unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds. Thanks to its task-generalization-oriented system design, H3 already possesses broad multimodal context understanding and generation capabilities at the pre-training stage, enabling outstanding performance in following complex multimodal instructions.
H3 supports the following input and output specifications:
Category Specification Output duration 4–15 seconds Output aspect ratio Supports a wide range of aspect ratios, including but not limited to 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16 Output resolution Supports various resolution dimensions. The shorter side is set to 768 pixels by default. 2K | generation can be achieved with H3-Regenerate-2K Output frame rate 24 FPS Output audio 32 kHz stereo Supported dialogue languages Stable support for 11 languages: Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, and Spanish. Additional languages are also supported to varying degrees



