Talking photo · 2026-08-04 · 12 min

Keyword: talking photo AI

How to Make a Talking Photo with AI Lip Sync

Make a talking photo AI clip from one portrait plus a script or recording. Visemes, crops, credits, and when to switch to an avatar instead.

  1. What a talking photo actually is
  2. Talking photo vs avatar, video lip sync, and pets
  3. How to generate a talking photo in the studio
  4. Quality checklist before you export
  5. Credits for a typical talking photo
  6. Rights: you must own the face and the audio
  7. Common talking-photo failures

What a talking photo actually is

A talking photo is a still with a visible mouth, driven by speech. You are not training a reusable host. You are making this picture talk: the founder, a customer who signed a release, a mascot, a family portrait. LipSyncing AI scores the audio into visemes and rewrites the mouth pass so the portrait stays itself instead of drifting into a cousin.

People search talking photo AI, AI talking photo generator, and make a photo talk for the same job. Category tools such as Magic Hour, Morphed, and Creatify sell photo-to-talking-video with different defaults. On this site the dedicated page is Talking Photo. Photo Lip Sync is the same engine with a photo-first layout. Talking Head Video is the presenter framing of that still.

Talking photo vs avatar, video lip sync, and pets

Use Talking Photo when the pixels of this still matter: a campaign portrait, a testimonial, a product designer in a real office. Use the AI Avatar Generator when you will reuse the same identity with new scripts every week. Mixing those jobs is how a channel looks like a different relative every Thursday.

Use Video Lip Sync when you already shot a take and only need the mouth rewritten. Use Talking Dog, Talking Cat, Anime Lip Sync, or Cartoon Lip Sync when the mouth is not a human cupid's bow. A people model will melt a snout or a seven-pixel anime line. Pick the specialist page; do not hope a slider saves it.

How to generate a talking photo in the studio

Open Talking Photo or the homepage studio. The controls are the same: model, resolution, aspect, duration, then a face and a line. There is no timeline. You are not keyframing a jaw. You are handing the engine a still and audio and letting it write visemes.

Keep the first version short. A talking photo that runs past thirty seconds asks a single still to invent a lot of life. If the script is a landing-page pitch, split it into two clips instead of padding silence to hit a minute.

  • Crop a front-facing still: even light, mouth visible, one main face, no microphone covering the lips.
  • Set 9:16 for TikTok or Reels, 16:9 for a site hero, 1:1 if the placement is a square ad.
  • Type a script and pick a voice (Nova, Rowan, Mira, Kai, Sol, or Lumen), or drop an audio file, or record in the browser.
  • Match duration to the line. Do not leave two seconds of polite nothing at the start — that still spends the 150-credit first block.
  • Press Make a video. Watch the local preview before you spend credits on an MP4.
  • If a plosive misses, shorten the script or replace noisy audio and generate again. Do not raise resolution to fix a late mouth.

Quality checklist before you export

A talking photo fails in the source more often than in the model. If the camera cannot see a mouth, the engine cannot honest-to-goodness sync one. Proof the still the way you would proof a headshot for a press kit.

If any item on that list fails, do not spend the MP4. Recrop, relight, or replace the audio and preview again. A talking photo that almost works usually needs a better still, not a higher resolution.

  • Both eye corners and both lip corners are in frame for the whole crop.
  • Teeth or a clear lip line are visible; lipstick or beard is fine, a scarf over the mouth is not.
  • Background is stable. A busy poster behind the head is fine; a second face in the same size is not.
  • Audio is dry. Music beds and restaurant noise smear phoneme timing.
  • The first viseme lands on the first syllable, not after a fake smile hold.
  • Head motion is a few degrees, not a nod that looks like a bobblehead.

Credits for a typical talking photo

One credit equals $0.01 on Starter, Pro, and Ultra. The first 30 seconds of any generation cost 150 credits ($1.50). Every extra second costs 5 credits ($0.05). A 45-second clip is 225 credits. A 60-second clip is 300 credits. Duration drives the meter. Resolution, aspect, and language do not.

Most talking photos should stay inside the first 30-second block. A 12-second greeting and a 28-second explainer both cost 150 credits on a plan. Stretching a thin still to 45 seconds costs 225 credits and usually looks worse, not better.

There is no free plan. The studio can still play a local preview on this computer before you pay. Finished MP4s need Starter, Pro, Ultra, or a one-time pack. Packs are $0.08 per credit and never expire. Yearly is 30% off the monthly price, paid once for 12 months, and does not auto-renew. Subscription credits spend first, then pack credits. Unused plan credits roll over while the subscription is active and cancel when you cancel.

Rights: you must own the face and the audio

Commercial use is allowed on paid plans when you have rights to the face and the audio. That is a legal fact, not a vibe. A celebrity crop from search, a stock model you did not license for AI video, or a voice you cloned without permission will get the clip taken down even if the visemes are perfect.

Preview does not change the rule. A local playback of a face you do not own is still a likeness you should not have uploaded. Paid plans allow commercial use of output you have rights to — they are not a license to other people's portraits.

  • Use your own portrait, a teammate who agreed in writing, or a generated / library look you are allowed to animate.
  • Do not animate customers, kids, or private photos from a camera roll that is not yours.
  • Scripts you write are yours. Licensed copy still needs the license.
  • Clone a voice only from a sample you own or have written permission to clone.
  • Pets and mascots follow the same rule: you need the right to use that animal or character on camera.

Common talking-photo failures

Most bad outputs are predictable. Fix the input instead of generating the same crop at 4K. Higher resolution does not invent a mouth the still does not contain.

When two failures stack — a profile crop and a noisy cafe recording — pick one to fix first. Start with the mouth visibility. Then trim the audio. Then generate. Buying a 60-second render of a bad crop is 300 credits of the same mistake.

  • Profile or three-quarter faces: the far lip corner is a guess. Shoot or crop closer to camera-facing.
  • Hands, food, or a mic in front of the lips: those frames cannot take visemes.
  • Group photos: pick one face and crop. Dual detection is not the talking-photo job.
  • Screenshots of screenshots: mush. Use the original file.
  • A paragraph of script on a passport crop: the still runs out of life. Split the line.
  • Pets and drawings run through this page: use Talking Dog, Talking Cat, or Anime Lip Sync.

Related posts