Sign In

Ace Step 1.5 XL Turbo and SFT - TEXT to AUDIO model with Ollama

Download

1 variant available

Type

Workflows

Stats

220

Reviews

Published

Apr 14, 2026

Base Model

Other

Hash

AutoV2
DD090FA526
Howling Aurora
tremolo28's Avatar

tremolo28

The Workflow was setup to have a clean "GUI" showing only parameters that matter, so you might want to toggle off Link visibility, like in above screenshot.


V1.6 Ace Step 1.5. Turbo and SFT normal and XL model with Ollama. Text to Audio/Song (examples below)

updated the settings for XL models and added a 3rd System Prompt for tags to chose, with more descriptive song descriptions.

1.5 XL SFT pipeline now has an "Adaptive Projected Guidance" node and negative prompt.

** See below some tips which model and settings to start with.


V1.5 Ace Step 1.5. Turbo and SFT normal and XL model with Ollama:

  • setup to create up to 4 tracks in a run, 2x Ace1.5 and 2x Ace1.5 XL, each with Turbo and SFT model, to compare (can be individually switched on/off)

  • VAE changed to tiled Audio VAE decode, uses less Vram.


V1.2 Ace Step 1.5. Turbo and SFT model with Ollama Text to Audio/Song

  • small update to GUI, system prompts and SFT sampler "engine"


V1.0 Ace Step 1.5. Turbo and SFT model with Ollama Text to Audio/Song

Ace Step uses TAGS and LYRICS to create a song. These can be generated by Ollama or by own prompts.

  • Can use any Song, Artist as reference or any other description to generate tags and lyrics.

  • Will output up to two songs, one generated by Turbo model, the other by the SFT model (experimental).

  • Keyscales, bpm and song duration can be randomized.

  • able to use dynamic prompts.

  • creates suitable songtitle and filenames with Ollama.

  • Lora Loader included, hope to see some Loras soon!

Avoid sage attention in your comfyui starting parameters, avoid --lowvram setting, as this might force Texencoder to run very slow on CPU instead of GPU.


Download Files:

Ollama Models, required for tags, lyrics and songtitle, you can choose 1,2 or 3 different models, tags and lyrics might need a bigger model >7b, songtitle can use a smaller model:


Alternative Turbo and SFT Models (normal, non XL) :


GGUF Models "normal" and XL: https://huggingface.co/Serveurperso/ACE-Step-1.5-GGUF/tree/main


Which models to start with ?

  • My current choice for normal model: Turbo-SFT merge_ta_0.5 & SFT-Shift1, using these settings:

    • Turbo-SFT_merge model with sampler: er_sde, scheduler: beta57 (or beta), 22 steps

    • SFT-Shift1 model with sampler euler, scheduler: normal, 138 steps

  • XL Model settings:

    • XL Turbo-SFT merge model: sampler: er_sde, scheduler: sgm_uniform, 40 steps

      • alternative: sampler: res_s2, scheduler beta57 (requires RES4LYF custom nodes)

    • XL SFT model: sampler: euler (or res_2s), scheduler: normal, 46 steps, CFG = 7.3, Adaptive Projected Guidance: eta = 1.05, norm_thresh= 1.3, momentum=0.0. Increase norm_thresh as the main parameter. These settings deliver "stabil" output for XL SFT,Base and their merges. The merges sound way better, pure SFT or Base introduce a lot of noise. I bypassed ModelSamplingAuraflow (see node next to model loader node). I think the base-turbo XL model merge fits well in that slot.

  • Disable "generate_audio_codes" in "TextEncodeAceStep" node to get different results, it works very well for many genres and reduces process time.

  • Ollama Model: Llama-3-NeuralDaredevil-8b-abliterated

More infos on models see thread below in discussion.


Save Location:

  • 📂 ComfyUI/

  • ├── 📂 models/

  • │ ├── 📂 diffusion_models/

  • │ │ └── acestep_v1.5_turbo.safetensors

  • │ ├── 📂 text_encoders/

  • │ │ ├── qwen_0.6b_ace15.safetensors

  • │ │ └── qwen_4b_ace15.safetensors (or 1.7b)

  • │ └── 📂 vae/

  • │ └── ace_1.5_vae.safetensors


Custom Nodes used:

optional (use Beta57 scheduler for a bit more punch, requires RES4LYF): https://github.com/ClownsharkBatwing/RES4LYF


Examples various styles:


Ollama help:

  1. Install Ollama from https://ollama.com/

  2. download a model: Go to a model page, chose a model , then hit the copy button, i.e. https://ollama.com/mirage335/Llama-3-NeuralDaredevil-8B-abliterated-virtuoso

  3. open terminal and paste the model name, i.e.: ollama run huihui_ai/qwen3-vl-abliterated

  4. model will be downloaded and can be selected in green comfy node "Ollama Connectivity". Hit "Reconnect" to refresh.