Product narration
Write a launch script in StepAudio 3 TTS and export a spoken take for demos and onboarding.
2 credits per request · reserved once before generation
Model: stepaudio-3-tts
Click a card to load the prompt into the studio, then review the credit cost before generating.
StepAudio 3 gives creators one workspace for StepFun audio. Start with text-to-speech on StepAudio 3 TTS, switch to ASR for recordings, talk live with Realtime, and create audio with Gen and Music. Credits are quoted before you submit.
Write a launch script in StepAudio 3 TTS and export a spoken take for demos and onboarding.
Drop a recording into ASR and keep an editable transcript instead of replaying the whole call.
Use Gen to mix speech, room tone, and cues into one clip for a storyboard or trailer.
Give Music a style caption and lyrics to hear a first chorus before a longer session.
One studio for speech, recognition, realtime agents, unified audio, and music — with credits shown before you submit.
Write a script, choose a system voice, and generate downloadable speech with an upfront credit quote.
Switch to ASR for recordings that should become editable transcripts instead of new audio.
The same studio hosts Realtime voice conversations with barge-in, unified audio generation, and music creation.
Signed-in generations stay in your account so you can replay, download, and iterate without losing the last good take.
Start on TTS by default, or switch mode when you need ASR, Realtime, Gen, or Music.
Paste text for speech, upload audio for transcription, or give a creative brief for Gen and Music.
Confirm the model and cost before submission. Credits are reserved before the provider request starts.
Listen in the browser, download the file, and find completed work in your account history.
See the credit cost before every generation, then choose a subscription or a one-time top-up when you need more.
Try StepAudio 3 TTS after sign-in.
For regular individual creative work.
For campaigns and higher-volume production.
For larger workloads, with adjustable credit capacity.
Capacity uses the lowest-cost live setup for this site. Your exact credit cost changes with the selected model and settings and is shown before generation.
StepAudio 3 is this site's product name for a StepFun audio studio covering TTS, ASR, Realtime, Gen, and Music. The homepage keyword and primary CTA center on StepAudio 3 TTS.
Text-to-speech uses StepFun model stepaudio-3-tts through the official audio speech API.
No. StepAudio 3 is an independent product that calls StepFun APIs. StepFun and its model names belong to their respective owner.
Generation stays disabled until API credentials, cost gates, and billing checks pass the launch SOP. The studio UI can still be reviewed while live generation is off.
StepAudio 3 is a focused studio for StepFun audio models. Generate speech with StepAudio 3 TTS, transcribe with ASR, try Realtime voice, create audio with Gen, or compose with Music.
Open the studio