Voice-first host · audio to avatar
Audio to Avatar
Audio to avatar is for when the performance already exists: a podcast clip, a lesson VO, a founder voice note, a localized dub. Do not run it through text-to-speech and lose the pacing. Upload the file, pick the face you intend to reuse, and generate. The mouth should follow this take — including pauses you meant to keep. If you only have a script and no recording, use Text to Avatar instead.
Three steps. That is the whole job.
- 1
Upload a video or photo
A clear, front-facing face works best. Or pick a sample in the studio.
- 2
Add audio or a script
Type the line and pick a voice, drop a file, or record. No timeline.
- 3
Generate lip sync
AI matches the mouth to every word. Then download the MP4.
Avatar video
Sync this vocal
150 credits · $1.50
Upload a video or photo
Pick a face from Character Library or upload a still, paste a script, then generate talking video.
Character library
Pick a saved look or a public character, then generate. You can still upload your own still.
No avatar selected yet
Public library
52 looks
Add audio or a script
Generate lip sync
150 credits · $1.50 · Log in
Press Make a video. If we do not have an API key yet, you still get a preview on this computer. AI by Zoice Avatar X
Your video
Nothing here yet. Add a face, add words, then press Make a video. The clip will play in this box.
See an example
Why people use audio to avatar
Keep the take you already like
Pronunciation, breath, and emphasis stay in the file. We are not rewriting the read.
Face is separate from the booth
The speaker on camera does not have to be the person who recorded — as long as you have rights to both.
Clean audio still wins
A dry vocal beats a café mix. Run Audio Cleaner first if the room is the problem.
Same credit meter as other talking jobs
150 credits for the first 30 seconds, 5 credits per extra second. Duration follows the file.
Three simple steps
- Step 1
Upload the vocal
WAV or a high-bitrate MP3. Trim silence you do not want to pay for.
- Step 2
Pick the avatar
Library look or a portrait with a visible mouth.
- Step 3
Generate and listen with the picture
If a plosive smears, recrop the face or recut the audio.
What this clip costs in credits
Every audio to avatar job uses the same meter: 150 credits ($1.50) for the first 30 seconds, then 5 credits ($0.05) per extra second. Subscriptions price credits at $0.01; one-time packs are $0.08 each.
| Duration | credits | Generation price |
|---|---|---|
| 15s | 150 | $1.50 |
| 30s | 150 | $1.50 |
| 45s | 225 | $2.25 |
| 60s | 300 | $3.00 |
| 90s | 450 | $4.50 |
| 120s | 600 | $6.00 |
First 30 seconds = 150 credits ($1.50). Plan credits roll over while you stay subscribed.
Good for
Lesson audio already in the LMS
Attach the chapter VO to a stable instructor face.
Podcast clips that need a face
A talking host on the quote without restaging the interview.
Localized dubs
A new language track on the same avatar, mouth rebuilt for that vocal.
Common questions about audio to avatar
You may also like
Each of these is a different job. Use the left menu, or tap one of these cards.
Text to Avatar
Turn a written script into a talking avatar video. Pick a library face or upload a portrait, generate speech, and LipSyncing AI matches the mouth to every line.
Talking Avatar
Build a talking avatar from a photo or a library character. Drive it with a script or audio and reuse the same host across every video.
Audio Cleaner
Clean a voice recording before lip sync. Upload a WAV or MP3; LipSyncing AI previews a denoise pass you can attach to a talking photo or avatar.
AI Voice Cloning
Clone a voice from a short sample. The clone is saved to Voice Library so you can hear it, generate speech on Text to Speech, or pick it in talking photo, avatars, and other tools.