Sign In

MiniMax Music 3 - ComfyUI Song Generator Workflow (Local AI Music with Vocals)

Download

1 variant available

Config Other

audio_minimax_music_3.json

41.59 KB

Verified:

Type
Workflows
Stats

94

Reviews
Published

Aug 17, 2026

Base Model

MiniMax H3

Hash
AutoV2
02B89728AD
default creator card background decoration
Followers - 115

115

Likes - 281

281

MiniMax H3 is licensed by MiniMax under the MiniMax H3 Community License Agreement. That agreement’s Applicable Territory excludes the European Union, the United Kingdom, the Republic of Korea and the United States of America. Your use of H3 and of any H3 derivative is subject to that agreement and its Acceptable Use Policy.

MiniMax H3

FREE MiniMax Music 3 ComfyUI Workflow & Guide - Local AI Song Generator.png

A ready-to-run ComfyUI workflow for MiniMax Music 3, the new open-weights AI music model — one of the best local Suno alternatives out right now. Feed it lyrics and a music description and it generates a complete song with vocals, instruments, and full structure, up to 5 minutes, in 32 kHz stereo. Runs locally, no subscription, and light enough for ~8GB VRAM territory.

The example workflow is attached to this post — drag the .json into ComfyUI and load your three models.

Required Models

Update ComfyUI to 0.33.0+ first. Three files from the Comfy-Org/MiniMax-Music-3 repo:

  • minimax_music3_dit_fp16.safetensorsmodels/unet/ (I recommend FP16 over the INT8 for quality; INT8 exists if you're tight on space)

  • minimax_music3_text_encoder_pruned_int8_convrot.safetensorsmodels/clip/

  • minimax_music3_dav.safetensorsmodels/vae/

If the model doesn't show in the node dropdowns, a file is in the wrong folder — that's the usual cause.

Writing the Caption

Output quality lives in the caption. Write it in three parts — Global metadata (genre, BPM, key, mood, production), Vocal details (gender, tone, delivery), and Arrangement (how sections evolve) — and tag your lyrics with [Verse] / [Chorus] on their own lines. Easiest approach right now is to have an LLM format your idea into that structure while keeping your lyrics intact.

Notes

  • FP16 for final renders — noticeably better than INT8 if you have the VRAM.

  • Keep first generations short to dial in the caption before committing to a full track.

  • Generation time scales with length and GPU.

  • Licensing: MiniMax Music 3 uses MiniMax's community license (not fully open). Free for personal use; if you monetize, you're asked to credit "MiniMax-Music3" and disclose AI-generated audio. Full terms on the HF page.


More From Me

If you're into local AI video too, check out my MiniMax H3 work — the sister model that does text/image to video with native audio:

More free workflows, one-click installers, and tutorials: ▶️ YouTube - https://www.youtube.com/@TheLocalLab 🎨 Patreon - https://www.patreon.com/TheLocalLab 🛒 https://www.locallabdigest.com

If this workflow helped, a ❤️ or review is appreciated!