Drop in a sermon transcript, a manuscript, an interview, a whole book. Voices strips the timestamps, finds every speaker, casts a different voice for each one, adds a cinematic soundscape, and turns it into a finished audiobook.
SRT/VTT cues, "Name:" labels, and interview formats are detected. Every speaker gets their own voice, pitch, and pace.
Cinematic score, rain, ocean, fireside, wind, deep space, heartbeat — synthesized live, auto-ducked under the voice.
Amplified, Cathedral, Radio, Telephone. Premium neural voices supported with your own ElevenLabs key.
Record the full mix in one click, or render true MP3/WAV chapters with premium voices. Chapters, playlists, the works.
Effects apply to the soundscape mix and to premium-voice renders. Built-in device voices play clean by design of the browser.
One click. Your browser will ask to share this tab with audio — allow it, and Voices performs the chapter and captures voice + soundscape into a downloadable audio file. Works best in Chrome or Edge on desktop.
Plug in your own ElevenLabs API key to render true studio-grade MP3 chapters with ultra-realistic neural voices, mixed with your soundscape. Your key stays in this browser only.