Beating AI News Flash: StepFun releases StepAudio 3, launching five voice models in one go: Realtime, ASR, TTS, Gen, and Music.
In two speech evaluations by Artificial Analysis, StepAudio 3 Realtime ranked first in both. It scored 98.9% on Conversational Dynamics, higher than Qwen Audio 3.0 Realtime Plus's 98.4% and GPT-Realtime-2 High's 95.3%; it scored 99.7% on Speech Reasoning, also ranking first. The former mainly tests whether a model can handle pauses, turn-taking, user interruptions, and acknowledgments like "uh-huh" and "right," while the latter tests whether a model can directly understand audio and complete reasoning.
ASR also achieved 1.7% WER (word error rate) on Artificial Analysis's non-streaming speech recognition leaderboard, tying for first place with Fun-Realtime-ASR-preview. In addition, StepAudio 3 Gen can generate vocals, sound effects, ambient sound, and music all at once, and orchestrate them uniformly according to the scene.
All five models have now entered StepFun's voice product line.

