Home

Lip sync · 2026-09-06 · 13 min

Keyword: video lip sync

Video Lip Sync vs Talking Photo: Pick the Right Workflow

Video lip sync vs talking photo: keep a filmed performance or invent life on a still. How to pick the workflow, steps, credits, and failure modes.

  1. Video lip sync keeps the take you already shot
  2. When to pick video, still, dubbing, or talking head
  3. How to resync a recorded take
  4. Quality checklist: performance preserved
  5. Credits for the length of the tape
  6. Rights on footage and replacement lines
  7. Common video-lip-sync failures

Video lip sync keeps the take you already shot

Video lip sync is for clips that already have performance in them — a presenter walking a product, a founder on a webcam, an actor in a medium shot. You keep every gesture. The model only rebuilds the mouth so a new line, a new language, or a cleaned-up take lands on the picture you already shot.

Talking Photo starts from a still and has to invent head and eye movement. That is the right call when there is no tape. It is the wrong call when the tape has blocking you paid for. Category tools often hide both behind one upload well. LipSyncing AI splits them so you do not accidentally flatten a performance into a laminated portrait.

When to pick video, still, dubbing, or talking head

Use Video Lip Sync when picture lock matters: wardrobe, hands, walk-and-talk. Use Talking Photo when all you have is a JPEG. Use Talking Head Video when you want a presenter series from a portrait, not from a filmed set. Use Video Dubbing when the language must change, not just the wording.

Use Motion Control when you need to copy a performance onto a different face, then lip sync if the line is new. Use Avatar Dialogue for two people you do not have on tape. If you are unsure, ask: would I be sad if the shoulders stopped moving? If yes, you are on video lip sync.

How to resync a recorded take

Stable face, new vocal, same blocking. That is the whole graph.

If the new line is much longer than the original mouth flap, the performance will look like stalling. Rewrite to the take length or cut picture. The model will not magically invent new blocking.

  • Upload the MP4/MOV with a face large enough to see lip corners.
  • Choose audio: rewritten script plus voice, a new WAV, or a cleaned original.
  • Match duration to the take. Do not pad a 20-second shot to 60.
  • Generate. Preview locally. Watch whether hands and walks survived.
  • If a syllable is early, trim the audio and run again — do not raise resolution.
  • For a language change, stop and use Video Dubbing or the translator instead of this page alone.

Quality checklist: performance preserved

Mute the output. You should still see the original acting. Unmute. The mouth should be the only obvious rewrite.

Scrub a gesture you remember from the shoot — a point, a shrug. If it vanished, you accidentally ran a still tool or a motion tool. Video lip sync should leave that shrug alone.

  • Shoulders, hands, and walks match the source.
  • Mouth closes on the new line's plosives.
  • No cousin face — identity stayed.
  • Brief turns are OK; long profiles look guessed.
  • Room tone is not doubled under a new VO unless you planned it.
  • Two-person scenes were split or cropped; overlapping speech was not dumped in one pass.

Credits for the length of the tape

You pay for output duration, not for how expensive the shoot was. One credit = $0.01 on a plan. First 30 seconds = 150 credits ($1.50). Extra second = 5 credits ($0.05). 45s = 225 credits. 60s = 300 credits. A 90-second webcam update is 450 credits.

Talking Photo on a 12-second greeting is also 150 credits — the still is not cheaper because it is a still. Preview locally. No free plan. Packs are $0.08 per credit and never expire. Yearly is 30% off, no auto-renew. Subscription credits spend first, then packs. Unused plan credits roll over while the plan is active and cancel if you cancel. Cut the tape before you generate.

Rights on footage and replacement lines

You need rights to the picture and the new audio. Replacing dialogue does not give you a new license to the faces on screen. Talent releases should cover ADR-style AI mouth replacement if this is commercial.

Employees on a webcam all-hands may not have cleared a public YouTube with AI-replaced dialogue. Internal tape plus a new marketing line is a new use. Ask.

  • Own the footage or have a license that allows this alteration.
  • Own the new vocal, TTS library voice, or a consented clone.
  • Do not resync a celebrity interview you do not own.
  • Music in the source may need to be ducked or replaced separately.
  • Paid-plan commercial use still requires those rights.

Common video-lip-sync failures

Using a still tool on a performance, covering the mouth, and generating the whole timeline of a montage.

Montages fail because most frames have no mouth. Extract the talking-head shot, resync that, and rebuild the edit. Paying 300 credits to viseme a product macro is a waste.

  • Talking Photo on a walk-and-talk: you flattened the take.
  • Food, cups, hands, mics on the lips: those frames cannot viseme.
  • Whip-pans and postage-stamp faces: recrop or reshoot a lockoff.
  • Two speakers, one generate: split.
  • Language change without dubbing: English jaw on a new track.
  • Uncut 3-minute file: 900 credits of material you will not post.

Related posts