StepAudio 3 Music
modelYour notes
StepFun's long-form music generation model (with ACE Studio), released through the StepFun API with a technical report on 11 September 2026. The design is discrete-then-continuous: a StepAudio Music Tokenizer represents audio as a 50 Hz stream from a single 65,536-entry codebook, a Mixture-of-Experts autoregressive planner first writes an intermediate arrangement in ABC notation ("ABC-CoT") so harmony, rhythm, and melody are explicit in the generation context, and a flow-matching diffusion Transformer predicts continuous VAE latents that a decoder renders to 48 kHz stereo audio. A progressive curriculum and supervised fine-tuning cover song and instrumental generation, accompaniment from dry vocals, and cover-song synthesis up to 5 minutes 30 seconds, with DPO on top. The report claims the highest AudioBox content-enjoyment, usefulness, and production-quality scores and the highest MuQ-MuLan similarity among the evaluated systems, competitive SongBench results, and a Quality Elo of 1105 on the preliminary Artificial Analysis Music Arena vocals leaderboard, behind Suno V5.5 and Mureka. No weights are released.