Studio notes · 2026-03-12 · 13 min

Keyword: how AI lip sync works

How AI Lip Sync Works (Visemes, Not Magic)

How AI lip sync works: phonemes collapse into visemes, photos invent modest motion, video keeps the performance, and pets need a different mouth model.

  1. Speech is a timeline of shapes
  2. When this science belongs on a sibling tool
  3. A practical viseme workflow
  4. Quality checklist: visemes you can trust
  5. Credits still follow duration, not “complexity”
  6. Rights sit outside the model
  7. What still breaks, even when you understand visemes

Speech is a timeline of shapes

A recording is pressure over time. Lip sync software turns that pressure into phonemes — the sounds English, Spanish, or Japanese are made of — with start and end timestamps. Those phonemes collapse into visemes, the mouth shapes a camera can see. Several sounds share a shape: P, B, and M all close the lips. That compression is why a cheap filter looks almost right and never actually right.

LipSyncing AI scores the audio, then drives a mouth rig on your still or clip. The original performance — a gesture, a walk, a blink already in the footage — stays. Only the mouth pass is rewritten. On a photo there is no original performance, so the model invents modest head and eye motion so the still does not look laminated. That is the same photo-to-talking-video idea Zoice Avatar X and other talking-photo tools sell. It is not a 3D actor. It is a timed mouth on pixels you already have.

When this science belongs on a sibling tool

The homepage generator is the general viseme engine for people. Video Lip Sync is the same science with a constraint: do not invent a new performance, only retarget the mouth. Talking Photo invents a little life because a JPEG has none.

Singing Photo changes the timing prior: held vowels sit on the beat instead of conversational consonants. Talking Dog and Talking Cat swap the mouth topology. Anime Lip Sync and Cartoon Lip Sync drop photoreal teeth. Dubbing adds a translation and a new vocal before the viseme pass. If you only remember one rule: run the detector that matches the mouth you uploaded.

A practical viseme workflow

You do not edit visemes by hand in this studio. You give the engine a clean face and a clean track. The quality of those two files is the quality of the mouth.

You will be tempted to add settings. Resolution, 4K, a longer duration. None of those invent lip corners. The workflow is crop, dry audio, short preview, then a paid master of the take you would post.

  • Start with a face the camera can read: front-on, even light, mouth not covered.
  • Give it audio with a single speaker and almost no bed. Trim leading silence.
  • Generate a short preview locally. Watch plosives and the first syllable.
  • If the mouth is late, the audio has padding or noise — fix the track, not the resolution.
  • If the mouth is mush, the face is too small or too profile. Recrop.
  • Only then spend credits on the length you will actually publish.

Quality checklist: visemes you can trust

You are looking for a mouth that is wrong in specific, fixable ways — not “uncanny” as a mood. Name the failure and change the input.

Watch at 1x on a phone. Desktop full-screen hides late mouths. If you have to squint to see a P close, the face is too small in the crop.

  • Bilabials close. If P/B/M stay open, phoneme timing is off or the lips are occluded.
  • Wide vowels open. If every syllable is a polite half-smile, the still may be too tight or the audio too quiet.
  • Blinks land on pauses, not on stressed words.
  • Video jobs keep shoulders and hands from the source clip.
  • Photo jobs drift a few degrees. A 20-degree nod is a bug, not style.
  • Specialist mouths stay on-model: no human dentition on a drawing, no cupid's bow on a beak.

Credits still follow duration, not “complexity”

Viseme density does not change the price. A dense plosive line and a slow vowel line cost the same if they last the same number of seconds. One credit equals $0.01 on a plan. The first 30 seconds are 150 credits ($1.50). Every extra second is 5 credits ($0.05). 45 seconds is 225 credits. 60 seconds is 300 credits.

That is why you trim silence. Two seconds of nothing after the 30-second mark is 10 credits of air. A 15-second line still costs the full 150-credit first block — there is no discount for being short once you generate. Preview locally, then pay for the take you would post.

Packs are $0.08 per credit and never expire. Yearly is 30% off and does not auto-renew. There is no free plan. Subscription credits spend first, then packs. The science is visemes. The bill is seconds.

Rights sit outside the model

A viseme engine will happily animate a face you do not own. That is not permission. You need the right to the likeness and the right to the audio, including any clone.

The engine does not know who is in the still. That is your job before you press generate. A technically perfect viseme pass on a scraped actor is still a takedown.

  • Own the still or have a written release that covers AI video.
  • Own the vocal, a library voice, or a clone built from a sample you are allowed to use.
  • Do not scrape actors, streamers, or classmates “to see if it works.”
  • Copyrighted cartoon characters are not fair game because the mouth is small.
  • Commercial output on paid plans still requires those rights. The plan is not a license to other people's faces.

What still breaks, even when you understand visemes

Profiles, hands on lips, food, microphones covering the mouth, and faces smaller than a postage stamp. If the camera cannot see a mouth, the model cannot sync one. Crop first. Light the face. Trim silence so the first viseme is not two seconds of polite nothing.

Overlapping speakers confuse phoneme alignment. Crosstalk looks like a mouth chewing two conversations. Split the take. Extreme wide shots starve the detector of lip corners. Singing with a live band in the same file fights the pitch track — give the engine a dry vocal.

  • Late mouth: trim leading silence or replace a noisy file.
  • Melted pet or anime: you used the people generator.
  • English jaw on a dubbed clip: you skipped the dubbing / translator pass.
  • Bobblehead still: the photo job invented too much motion — recrop tighter on the face and shorten the line.
  • Perfect visemes on a stolen celebrity: that is a rights failure, not a quality win.

Related posts