ACE-Step: A step towards music generation foundation model

https://news.ycombinator.com/rss Hits: 15
Summary

A Step Towards Music Generation Foundation Model Project | Checkpoints | Space Demo | Discord Table of Contents 📢 News and Updates 🚀 2025.05.06: Open source demo code and model Release training code 🔥 Release training code 🔥 Release LoRA training code 🔥 Release LoRA training code 🔥 Release RapMachine lora 🎤 Release RapMachine lora 🎤 Release ControlNet training code 🔥 Release ControlNet training code 🔥 Release Singing2Accompaniment controlnet 🎮 Release Singing2Accompaniment controlnet 🎮 Release evaluation performance and technical report 📄 🏗️ Architecture 📝 Abstract We introduce ACE-Step, a novel open-source foundation model for music generation that overcomes key limitations of existing approaches and achieves state-of-the-art performance through a holistic architectural design. Current methods face inherent trade-offs between generation speed, musical coherence, and controllability. For instance, LLM-based models (e.g., Yue, SongGen) excel at lyric alignment but suffer from slow inference and structural artifacts. Diffusion models (e.g., DiffRhythm), on the other hand, enable faster synthesis but often lack long-range structural coherence. ACE-Step bridges this gap by integrating diffusion-based generation with Sana’s Deep Compression AutoEncoder (DCAE) and a lightweight linear transformer. It further leverages MERT and m-hubert to align semantic representations (REPA) during training, enabling rapid convergence. As a result, our model synthesizes up to 4 minutes of music in just 20 seconds on an A100 GPU—15× faster than LLM-based baselines—while achieving superior musical coherence and lyric alignment across melody, harmony, and rhythm metrics. Moreover, ACE-Step preserves fine-grained acoustic details, enabling advanced control mechanisms such as voice cloning, lyric editing, remixing, and track generation (e.g., lyric2vocal, singing2accompaniment). Rather than building yet another end-to-end text-to-music pipeline, our vision is to establish a foundation model f...

First seen: 2025-05-06 22:02

Last seen: 2025-05-07 12:04