StepFun's long-form music generation model (with ACE Studio), released through the StepFun API with a technical report on 11 September 2026. The design is discrete-then-continuous: a StepAudio Music Tokenizer represents audio as a 50 Hz stream from a single 65,536-entry codebook, a Mixture-of-Experts autoregressive planner first writes an intermediate arrangement in ABC notation ("ABC-CoT") so harmony, rhythm, and melody are explicit in the generation context, and a flow-matching diffusion Transformer predicts continuous VAE latents that a decoder renders to 48 kHz stereo audio. A progressive curriculum and supervised fine-tuning cover song and instrumental generation, accompaniment from dry vocals, and cover-song synthesis up to 5 minutes 30 seconds, with DPO on top. The report claims the highest AudioBox content-enjoyment, usefulness, and production-quality scores and the highest MuQ-MuLan similarity among the evaluated systems, competitive SongBench results, and a Quality Elo of 1105 on the preliminary Artificial Analysis Music Arena vocals leaderboard, behind Suno V5.5 and Mureka. No weights are released.

Model Details

Architecture MOE
License Proprietary (API)

Paper

audiogenerationmultimodalproprietary

Related