Speech synthesis for real-time voice agents
Hear the voice in your words
Try a line your application would say. Compare voices using the same words, then listen for pronunciation and pacing.
Build speech into your application
Sonic turns the text your application produces into audio. Your application supplies the response; the voice model speaks it.
Choose a voice
Test voices with the words your application will actually say. Include names, numbers, and abbreviations in your evaluation instead of relying only on a greeting.
Generate audio through the API
Send text from your application and receive synthesized audio. Choose a voice and output format, then play the response or save it as a file.
Stream spoken responses
Use streaming audio for conversational playback. Measure the time until users hear speech in your application, including network transfer and playback buffering.
Browser playback or a speech API?
Both turn text into speech. Choose based on where you want the voice to run and how your application needs to use the audio.
Browser speech synthesis
Use the browser's SpeechSynthesis interface for playback with device voices. The available voices depend on the browser and operating system.
Best for reading text aloud in a web page.
Read the browser TTS guideCartesia's speech API
Choose a Sonic voice by ID and generate speech from your backend. Stream, save, or play the audio in web, mobile, and phone applications.
Best for generating audio for your application.
Explore the speech APILifelike, expressive voices for every use case
Listen to speech samples for different applications. Compare the delivery, then try a voice with the words your audience will hear.
Support
Turn customer support answers into spoken responses with speech synthesis. Read service updates and troubleshooting steps in a consistent voice across calls and applications.
Gaming
Generate character dialogue from your game's scripts. Compare voices for each role, then listen to new lines in context to check pacing and delivery.
Content
Create voiceovers for videos, tutorials, and product explainers from text. Revise a script and generate new narration without recording each line again.
Media
Turn written articles and stories into spoken audio. Use text-to-speech for podcast introductions or narrated news, checking names and pronunciation before publishing.
Healthcare
Read appointment reminders and patient service information from your team's reviewed text. Use a consistent voice for routine instructions and scheduling updates.
Sales
Add spoken narration to product demonstrations and sales explainers. Test your script in different voices, then update the audio when your product or message changes.
Voice Agents
Give AI voice agents spoken responses generated from your application's text. Choose a voice for greetings and conversations, then test how it sounds across short and longer replies.
Dubbing
Generate speech from translated scripts for your localization workflow. Compare voices in supported languages and review pronunciation and timing before adding the audio to a video.
Avatars
Pair a digital avatar with synthesized speech for presentations and guided experiences. Generate dialogue from text and coordinate audio playback with your avatar's animation.
Logistics
Turn shipment updates and dispatch instructions into spoken messages. Connect speech synthesis to your application so the audio reflects the latest delivery information.
Recruiting
Create spoken interview introductions and candidate scheduling messages from approved scripts. Keep instructions consistent and regenerate the audio when your hiring process changes.
Accessibility
Offer audio versions of written guides, articles, and learning materials. Let people listen to your content and follow along with the original text at their own pace.
Fluent and native, worldwide
Reach international markets with Sonic — 44 languages and a wide range of accents, all with native-speaker quality voices.
From a voice demo to a conversation
Speech synthesis is the spoken response. A conversational agent also needs to understand input, decide what to say, and handle interruptions.
Build your audio pipeline
Use the Bytes API when your transcript is ready, or a WebSocket when text arrives in chunks. Match the output format to your player and stop pending audio when a user interrupts.
Read the WebSocket guideBuild with Managed Agents
Connect a voice with your agent's instructions and tools. Use Managed Agents when you want to work on the conversation alongside its voice.
Explore Managed AgentsSpeech synthesis FAQs
Get started today
Talk to an expert.
Connect with a member of our team and learn how Cartesia can help you build world-class voice experiences.
Start building.
Access our models via API and bring a voice agent into production in minutes.