AI Audio · AI text to speech
AI Text to Speech
Text to speech here is powered by ElevenLabs, Cartesia, and Chatterbox on RunPod. This page does not clone a voice. It lists the Voice Library for the engine you pick: My voices first, public voice AI second. Choose one, paste the script, generate the track, then send it to an avatar.
Text to speech
Generate this voiceover
150 credits · $1.50
Pick a voice from the library, paste a script, then generate. My voices are clones and designs you saved. Public voice AI is the catalog for this engine.
My voices0
No saved voices for this engine yet. They appear here after you clone or design one. You can still pick a public voice below.
Public voice AI10
Staff-curated public catalog for this engine, plus live voices when a key is pasted. Play a sample here. Open Text to Speech when you want a generated track.
Showing the built-in catalog. Staff can edit it on Admin → Voice Library.
Text to speech runs on ElevenLabs, Cartesia, and Chatterbox (RunPod). Pick a voice from My voices or public voice AI. Paste keys on Admin → Models.
See an example
Why people use AI text to speech
My voices, then public AI
Clones and designed voices you saved sit first. Public catalog voices sit second. Pick one and generate.
Emotion without a second take
Calm, bright, or urgent — match the landing page, not a flat read.
Made to hand off
The file is an input for talking photo, talking avatar, dubbing, and singing (spoken intros).
Script-length honesty
TTS is cheap compared with a crew. The video job still bills the 30-second block plus extras.
Three simple steps
- Step 1
Pick a voice from the library
My voices are clones and designs you already saved. Public voice AI is the catalog.
- Step 2
Paste the script and generate
Listen before you spend a talking-video job.
- Step 3
Send to a face tool
Talking avatar, photo lip sync, or video lip sync.
What this clip costs in credits
Every AI text to speech job uses the same meter: 150 credits ($1.50) for the first 30 seconds, then 5 credits ($0.05) per extra second. Subscriptions price credits at $0.01; one-time packs are $0.08 each.
| Duration | credits | Generation price |
|---|---|---|
| 15s | 150 | $1.50 |
| 30s | 150 | $1.50 |
| 45s | 225 | $2.25 |
| 60s | 300 | $3.00 |
| 90s | 450 | $4.50 |
| 120s | 600 | $6.00 |
First 30 seconds = 150 credits ($1.50). Plan credits roll over while you stay subscribed.
Good for
Avatar reads
Generate audio, then open talking avatar on the same script.
Dubbing drafts
TTS in the target language before you hire a human mixer.
IVR-adjacent help clips
Consistent library voice across 40 macros.
Common questions about AI text to speech
You may also like
Each of these is a different job. Use the left menu, or tap one of these cards.
AI Voice Cloning
Clone a voice from a short sample. The clone is saved to Voice Library so you can hear it, generate speech on Text to Speech, or pick it in talking photo, avatars, and other tools.
AI Voice Design
Describe the voice, pick ElevenLabs or Cartesia, listen to the preview, then save it to Voice Library or discard it.
Voice Library
Browse voices you cloned, then public voice AI: ElevenLabs premade and shared voices, Cartesia’s catalog, and Chatterbox presets. Listen to samples here. Generate a script on Text to Speech.
Talking Avatar
Build a talking avatar from a photo or a library character. Drive it with a script or audio and reuse the same host across every video.