Audio · 2026-09-02 · 12 min
Keyword: text to speech lip sync
Text to Speech + Lip Sync: From Script to Talking Video
Text to speech lip sync: type a script, pick a voice or clone, generate visemes. How TTS, voice cloning, and talking video fit in one studio.
- TTS is the track the mouth will follow
- TTS vs cloning vs recording vs dubbing
- From script to talking video, in order
- Quality checklist for speech that visemes
- Credits for audio and for the mouth
- Rights on voices
- Common TTS plus lip-sync failures
TTS is the track the mouth will follow
Text to speech lip sync is two jobs in order. First you turn a script into a vocal. Then you match a mouth to that vocal. LipSyncing AI keeps both in one studio so you are not exporting a WAV from one site and hoping visemes line up on another. This page is the audio half. Talking Photo, Talking Avatar, Video Lip Sync, and Spokesperson are the face half.
Library voices (Nova, Rowan, Mira, Kai, Sol, Lumen, plus the wider catalog on the TTS tool) are for speed. Voice cloning is for a specific throat you have the right to copy. Do not clone a stranger. Do not expect TTS to be a song — sung jobs want a real vocal on Singing Photo.
TTS vs cloning vs recording vs dubbing
Use Text to Speech when the voice can be a house narrator. Use AI Voice Cloning when it must be a specific person you have permission to clone, then generate the new script in that voice. Use in-browser recording when the take should be you, today, with the mic you have. Use Video Dubbing or AI Video Translator when the picture already exists in another language.
Skip TTS when you already have a great WAV. Drop the file on the talking tool instead. Skip a talking generate when you only needed a voiceover on B-roll with no face — TTS alone is cheaper in time even though the duration meter is the same, because you will not run a second mouth pass.
From script to talking video, in order
Write for the ear. TTS will read the commas you forgot and the URLs you should have spelled out.
If the talking tool can take the script and the same voice directly, use that and skip a separate billed TTS export. Double-running “to be safe” is how a 25-second line becomes two 150-credit first-blocks.
- Paste the script. Estimate duration from ~150 words per minute; the studio caps at 180 seconds.
- Pick a library voice, or generate from a clone you set up on Voice Cloning.
- Preview the audio. Fix names, numbers, and CTAs before you attach a face.
- Open Talking Photo, Talking Avatar, Spokesperson, or Video Lip Sync with that audio (or paste the same script there and pick the same voice).
- Generate the mouth pass. Watch plosives on the product name.
- If the mouth is late, the TTS file has leading silence — trim and rerun the face job, not the whole essay.
Quality checklist for speech that visemes
If you would not put this VO on a podcast, do not put it on a face. The mouth cannot save a muddy read.
Listen on a phone speaker. If names collapse, respell them phonetically in the script. The mouth will follow the sound you actually generated, not the spelling you meant.
- Numbers and SKUs are pronounced as you intended.
- No music bed under the TTS file you send to lip sync.
- Sentences are short enough that visemes have somewhere to land.
- Voice energy matches the on-screen person.
- Clone quality was checked on a 10-second sample before a 60-second spend.
- The talking export is a separate proof: ears, then eyes.
Credits for audio and for the mouth
Speech jobs and talking jobs both follow duration. One credit = $0.01 on a plan. First 30 seconds = 150 credits ($1.50). Extra second = 5 credits ($0.05). 45s = 225 credits. 60s = 300 credits. A 25-second TTS file plus a 25-second talking photo is two 150-credit first-blocks if you generate both as billed jobs.
When the talking tool can take the script directly, you may only pay once for the combined generate — do not double-run TTS and talking “to be safe” without checking. Preview locally. No free plan. Packs are $0.08 per credit and never expire. Yearly is 30% off, no auto-renew. Subscription credits spend first, then packs. Unused plan credits roll over while the plan is active and cancel if you cancel.
Rights on voices
Library voices are licensed for the studio. Clones are not a loophole. Only voices you own or have written permission to clone. Impersonation is not allowed. The face on the later talking job has its own rights, the same as any talking photo.
A clone of your own voice is the cleanest commercial path. A clone of a contractor needs a contract that says AI synthesis, ads, and the locales you will ship. Handshake permission in a Slack thread is a weak place to stand.
- Do not clone YouTubers, actors, or coworkers “as a test.”
- CEO clones need the CEO.
- Scripts with unlicensed third-party copy are still a problem even in a stock voice.
- Recording in the browser is yours if it is your voice.
- Commercial use on paid plans still requires those permissions.
Common TTS plus lip-sync failures
Bad pronunciation, double billing, and clones you should not have made. Two of those are workflow. One is a stop.
Mismatched script and WAV is the silent killer: you edited the text after you generated audio and then wondered why visemes missed. One source of truth. Then the face job.
- URL read as a slur of letters: spell it how it should sound.
- Singing requested from TTS: record a vocal or use a sung file on Singing Photo.
- Noisy clone sample: 10 clean seconds beat an hour of restaurant audio.
- Mouth pass on a different script than the WAV: they will never match.
- Dubbing a tape with TTS on the homepage: use Video Dubbing or the translator.
- 60-second retakes of a 12-second line: trim, preview, then pay.
Related posts
AI voice cloning lip sync
Voice Cloning for Lip Sync: A Voice You Own, a Mouth That Matches
AI voice cloning lip sync: clone a voice you own, generate the script, then match visemes so the mouth follows a throat that is actually yours.
AI talking head generator
AI Talking Head Generator: Photo to Presenter in One Pass
An AI talking head generator that turns a photo into a presenter: one pass from portrait and script to a lip-synced talking head video.
AI spokesperson video
How to Make an AI Spokesperson Video for a Landing Page
How to make an AI spokesperson video for a landing page: a professional presenter, a proofed script, lip sync, and credits for a 30-second read.