Local Ai
ace step 1.5 xl sft terrible results
A Reddit thread on r/StableDiffusion where a user reports poor output quality when using the ACE-Step 1.5 XL SFT model variant for AI music generation. The SFT (Supervised Fine-Tuning) variant of ACE-
A Reddit thread on r/StableDiffusion where a user reports poor output quality when using the ACE-Step 1.5 XL SFT model variant for AI music generation. The SFT (Supervised Fine-Tuning) variant of ACE-Step 1.5 XL is designed to be tuned for quality and variation, while the XL series uses a larger 4B-parameter DiT decoder requiring significant VRAM (≥12GB with offloading/quantization), which may contribute to configuration or hardware-related issues for some users. The discussion likely covers troubleshooting tips, settings comparisons, and community workarounds for getting better results from the model.
Related
- Echo Chamber - AceStep 1.5 song (XL version)
- ACE-Step 1.5 XL Base — BF16 version (converted from FP32)
- Musicvideo on local Hardware
- LTX 2.3 Lip Sync Music Clip -- Drake - Toosie Slide
Source: r/StableDiffusion | 2026-04-11