Script-first host · text to avatar
Text to Avatar
Text to avatar is the path when the copy is already approved and nobody wants another recording day. Paste the script, pick a reusable face from Character Library or a portrait you own, choose a voice, and generate. This is not cinematic text-to-video. It is a presenter who reads the words you typed, with visemes rebuilt on that face. Change a sentence, regenerate the chapter, keep the same host.
Three steps. That is the whole job.
- 1
Upload a video or photo
A clear, front-facing face works best. Or pick a sample in the studio.
- 2
Add audio or a script
Type the line and pick a voice, drop a file, or record. No timeline.
- 3
Generate lip sync
AI matches the mouth to every word. Then download the MP4.
Avatar video
Generate from this script
150 credits · $1.50
Upload a video or photo
Pick a face from Character Library or upload a still, paste a script, then generate talking video.
Character library
Pick a saved look or a public character, then generate. You can still upload your own still.
No avatar selected yet
Public library
52 looks
Add audio or a script
Voice library
Same catalog as Voice Library: your clones first, then public voice AI for the engine you pick.
Rachel
No clones for this engine yet. Clone or design one, or pick public voice AI. Clone a voice
Public voice AI
10 voices
Generate lip sync
150 credits · $1.50 · Log in
Press Make a video. If we do not have an API key yet, you still get a preview on this computer. AI by Zoice Avatar X
Your video
Nothing here yet. Add a face, add words, then press Make a video. The clip will play in this box.
See an example
Why people use text to avatar
The script is the source of truth
Edits happen in the text. You do not restage a booth because legal changed one product name.
Library face or your still
A saved avatar keeps the channel consistent. A portrait is fine when the host is a specific person you have rights to.
Voice attached, not implied
Pick a library voice or a clone. The mouth follows that take, not a silent prompt.
Not a scene generator
If you want landscapes and camera moves, use text to video. This page is a talking host.
Three simple steps
- Step 1
Lock the words
Short sentences. Spell names the way they should be heard.
- Step 2
Pick the face and a voice
Library look or an owned portrait, then TTS or a clone.
- Step 3
Generate the talking clip
Watch the first greeting. Fix the script, not the camera.
What this clip costs in credits
Every text to avatar job uses the same meter: 150 credits ($1.50) for the first 30 seconds, then 5 credits ($0.05) per extra second. Subscriptions price credits at $0.01; one-time packs are $0.08 each.
| Duration | credits | Generation price |
|---|---|---|
| 15s | 150 | $1.50 |
| 30s | 150 | $1.50 |
| 45s | 225 | $2.25 |
| 60s | 300 | $3.00 |
| 90s | 450 | $4.50 |
| 120s | 600 | $6.00 |
First 30 seconds = 150 credits ($1.50). Plan credits roll over while you stay subscribed.
Good for
Lesson scripts
A module outline becomes a talking chapter without filming the instructor again.
Product copy that has to speak
Landing-page paragraphs turned into a face-led explainer.
Help macros
Approved answers read by the same avatar every time the article updates.
Common questions about text to avatar
You may also like
Each of these is a different job. Use the left menu, or tap one of these cards.
Audio to Avatar
Sync an avatar to existing audio. Upload a voice take, pick a library face or portrait, and LipSyncing AI rebuilds the mouth to that recording.
Talking Avatar
Build a talking avatar from a photo or a library character. Drive it with a script or audio and reuse the same host across every video.
AI Text to Speech
Turn a script into speech with ElevenLabs, Cartesia, or Chatterbox. Pick a voice from My voices — clones and designs you saved — or from public voice AI, then generate.
AI Presenter
Create presenter-style talking videos from a portrait, a script, or a voice track. Built for lessons, product demos, onboarding, and explainers on LipSyncing AI.