StepAudio 3 Realtime
modelYour notes
StepFun's audio-language foundation model for spoken interaction, organised as a continuous listen–converse–think–act loop. Deep Perception reads acoustic cues for intent, Seamless Duplex models synchronised audio streams to handle pauses, backchannels, and interruptions, and Think-While-Speaking runs private reasoning in parallel with spoken delivery so deliberation does not cost latency. The 90-author report gives 73.0 macro-average on StepAudioChat in reasoning mode, 90.6 on MMSU, 98.9 on the Artificial Analysis Full-Duplex Bench, and 56.0% on τ-Voice. A companion report the day before describes StepAudio 3 Gen, a discrete autoregressive generator over 12.5 Hz residual-vector-quantised tokens (a shared 16 × 2048 code space) that covers zero-shot TTS, voice design, vocals, sound effects, music, and mixtures in one model, a departure from the diffusion-Transformer generation now common in audio. Neither report releases weights; both models are served through the StepFun API. Filed with StepAudio 3 Music as the third report of the series.