Download
1 variant available
This checkpoint includes a config file, download and place it along side the checkpoint.
940 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
(7)
Aug 17, 2026
MiniMax H3

1150 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
2810 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
MiniMax H3 is licensed by MiniMax under the MiniMax H3 Community License Agreement. That agreement’s Applicable Territory excludes the European Union, the United Kingdom, the Republic of Korea and the United States of America. Your use of H3 and of any H3 derivative is subject to that agreement and its Acceptable Use Policy.
MiniMax H3
A ready-to-run ComfyUI workflow for MiniMax Music 3, the new open-weights AI music model — one of the best local Suno alternatives out right now. Feed it lyrics and a music description and it generates a complete song with vocals, instruments, and full structure, up to 5 minutes, in 32 kHz stereo. Runs locally, no subscription, and light enough for ~8GB VRAM territory.
The example workflow is attached to this post — drag the .json into ComfyUI and load your three models.
Required Models
Update ComfyUI to 0.33.0+ first. Three files from the Comfy-Org/MiniMax-Music-3 repo:
minimax_music3_dit_fp16.safetensors →
models/unet/(I recommend FP16 over the INT8 for quality; INT8 exists if you're tight on space)minimax_music3_text_encoder_pruned_int8_convrot.safetensors →
models/clip/minimax_music3_dav.safetensors →
models/vae/
If the model doesn't show in the node dropdowns, a file is in the wrong folder — that's the usual cause.
Writing the Caption
Output quality lives in the caption. Write it in three parts — Global metadata (genre, BPM, key, mood, production), Vocal details (gender, tone, delivery), and Arrangement (how sections evolve) — and tag your lyrics with [Verse] / [Chorus] on their own lines. Easiest approach right now is to have an LLM format your idea into that structure while keeping your lyrics intact.
Notes
FP16 for final renders — noticeably better than INT8 if you have the VRAM.
Keep first generations short to dial in the caption before committing to a full track.
Generation time scales with length and GPU.
Licensing: MiniMax Music 3 uses MiniMax's community license (not fully open). Free for personal use; if you monetize, you're asked to credit "MiniMax-Music3" and disclose AI-generated audio. Full terms on the HF page.
More From Me
If you're into local AI video too, check out my MiniMax H3 work — the sister model that does text/image to video with native audio:
MiniMax H3 ComfyUI Workflow + Turbo LoRAs (4-8 step speedup, native audio) — on my Patreon: https://www.patreon.com/TheLocalLab
MiniMax H3 RunPod Template (run it in the cloud, no local GPU): https://get.runpod.io/Minimax-H3-ComfyUI
More free workflows, one-click installers, and tutorials: ▶️ YouTube - https://www.youtube.com/@TheLocalLab 🎨 Patreon - https://www.patreon.com/TheLocalLab 🛒 https://www.locallabdigest.com
If this workflow helped, a ❤️ or review is appreciated!


