Speech synthesis for real-time voice agents

Turn text into spoken audio with Sonic, Cartesia's text-to-speech model. Try a voice in the Playground, then generate audio from your application through the API.

Hear the voice in your words

Try a line your application would say. Compare voices using the same words, then listen for pronunciation and pacing.

242/500

Build speech into your application

Sonic turns the text your application produces into audio. Your application supplies the response; the voice model speaks it.

Choose a voice

Test voices with the words your application will actually say. Include names, numbers, and abbreviations in your evaluation instead of relying only on a greeting.

Generate audio through the API

Send text from your application and receive synthesized audio. Choose a voice and output format, then play the response or save it as a file.

Stream spoken responses

Use streaming audio for conversational playback. Measure the time until users hear speech in your application, including network transfer and playback buffering.

Browser playback or a speech API?

Both turn text into speech. Choose based on where you want the voice to run and how your application needs to use the audio.

Browser speech synthesis

Use the browser's SpeechSynthesis interface for playback with device voices. The available voices depend on the browser and operating system.

Best for reading text aloud in a web page.

Read the browser TTS guide

Cartesia's speech API

Choose a Sonic voice by ID and generate speech from your backend. Stream, save, or play the audio in web, mobile, and phone applications.

Best for generating audio for your application.

Explore the speech API

Lifelike, expressive voices for every use case

Listen to speech samples for different applications. Compare the delivery, then try a voice with the words your audience will hear.

Support

Turn customer support answers into spoken responses with speech synthesis. Read service updates and troubleshooting steps in a consistent voice across calls and applications.

Gaming

Generate character dialogue from your game's scripts. Compare voices for each role, then listen to new lines in context to check pacing and delivery.

Content

Create voiceovers for videos, tutorials, and product explainers from text. Revise a script and generate new narration without recording each line again.

Media

Turn written articles and stories into spoken audio. Use text-to-speech for podcast introductions or narrated news, checking names and pronunciation before publishing.

Healthcare

Read appointment reminders and patient service information from your team's reviewed text. Use a consistent voice for routine instructions and scheduling updates.

Sales

Add spoken narration to product demonstrations and sales explainers. Test your script in different voices, then update the audio when your product or message changes.

Voice Agents

Give AI voice agents spoken responses generated from your application's text. Choose a voice for greetings and conversations, then test how it sounds across short and longer replies.

Dubbing

Generate speech from translated scripts for your localization workflow. Compare voices in supported languages and review pronunciation and timing before adding the audio to a video.

Avatars

Pair a digital avatar with synthesized speech for presentations and guided experiences. Generate dialogue from text and coordinate audio playback with your avatar's animation.

Logistics

Turn shipment updates and dispatch instructions into spoken messages. Connect speech synthesis to your application so the audio reflects the latest delivery information.

Recruiting

Create spoken interview introductions and candidate scheduling messages from approved scripts. Keep instructions consistent and regenerate the audio when your hiring process changes.

Accessibility

Offer audio versions of written guides, articles, and learning materials. Let people listen to your content and follow along with the original text at their own pace.

Fluent and native, worldwide

Reach international markets with Sonic — 44 languages and a wide range of accents, all with native-speaker quality voices.

Most popular locales

From a voice demo to a conversation

Speech synthesis is the spoken response. A conversational agent also needs to understand input, decide what to say, and handle interruptions.

Build your audio pipeline

Use the Bytes API when your transcript is ready, or a WebSocket when text arrives in chunks. Match the output format to your player and stop pending audio when a user interrupts.

Read the WebSocket guide

Build with Managed Agents

Connect a voice with your agent's instructions and tools. Use Managed Agents when you want to work on the conversation alongside its voice.

Explore Managed Agents

Speech synthesis FAQs

Get started today

Talk to an expert.

Connect with a member of our team and learn how Cartesia can help you build world-class voice experiences.

Contact Sales

Start building.

Access our models via API and bring a voice agent into production in minutes.

Try Cartesia