Skip to main content
Guide7 min read

AI Talking Avatar: How to Make a Photo Talk

A photo and an audio file are the entire input. Here is how to get lip sync that holds up, and where the ethical line sits.

To make a photo talk with AI you need two files: a clear image of a face, and an audio recording of what it should say. The Talking Avatar model in VIBE takes those two inputs and returns a video of the person in the photo speaking your recording, with mouth movement matched to the audio. There is no text prompt, no rigging and no avatar to design — the recording sets both the words and the length of the video. It runs on iPhone and Android, outputs 480p or 720p, and costs a flat price per video rather than a rate per second.

What an AI talking avatar actually is

The name covers two very different things. One is the classic avatar platform, where you design a character, rig it, and drive it with a script — a lot of setup before you see anything. The other, and the one this article is about, is audio-driven photo animation: a model reads a single still image and an audio waveform and generates the frames in between, moving the mouth, jaw, cheeks and head to match the sound.

That is why the workflow is so short. Nothing is designed and nothing is animated by hand — the photo is the character. Change the photo and you have a different presenter; change the audio and the same presenter says something else.

It also explains the model’s two unusual properties. There is no prompt field, because the audio already specifies everything the model needs to do — no scene to describe, no camera move to direct. And there is no duration picker, because the video ends when the recording ends. A 22-second voice note becomes a 22-second video.

What you need before you start

Two files, and neither of them is hard to get.

  • A photo of a face. A portrait, a headshot, a photo of a colleague who has agreed to it, an illustration, or a character you generated with one of the AI image models in the app. It does not have to be a photograph of a real person.
  • An audio recording in MP3, WAV or FLAC. Your own voice recorded on your phone, a professional voiceover, a client-supplied track, or a synthetic voice generated elsewhere and saved as a file.

The audio does all the work, so it is the part worth thinking about. Write the script to be spoken rather than read: short sentences, one idea each, contractions where you would naturally use them. Copy that reads well on a page often sounds stilted out loud, and that is immediately obvious on a talking-head video where there is nothing else on screen to look at.

Step by step

  1. 1Record or export the audio first. It sets the length and the pacing of the whole video, so get it right before anything else, and trim the silence off both ends — dead air at the start becomes seconds of a face doing nothing.
  2. 2Pick the photo. Front-facing, evenly lit, face large enough that the mouth is clearly resolved. The output keeps the framing and background of whatever you choose, because there is no prompt to replace them.
  3. 3Open VIBE, choose the Talking Avatar model, and attach the photo and the audio file. There is no prompt box — the two files are the complete input.
  4. 4Generate a draft at 480p. It costs less than 720p and is enough to check the sync, the head movement, and whether the photo was a good choice at all.
  5. 5Watch it with the sound on, then with the sound off. With sound you are checking that the mouth matches the words; without it you are checking that the face looks natural on its own, which is where a bad source photo shows up.
  6. 6Regenerate at 720p once you are happy. Use 720p for a product page, an ad, or a client deliverable; 480p is usually fine for a social feed.

How to get clean lip sync

Lip sync quality is decided almost entirely by the two files you supply. There is no setting that rescues a bad input, so the fixes all happen before you generate.

Choosing the photo

  • Face the camera. A near-frontal angle shows both corners of the mouth. Strong three-quarter and profile angles force the model to infer geometry it cannot see, and the mouth movement gets vaguer as a result.
  • Keep the mouth unobstructed. Hands near the face, a microphone in front of the lips, a scarf, a beard covering the mouth line, or a hard shadow across the lower face all reduce sync accuracy.
  • Give the face some resolution. A small face in a low-resolution photo has very few pixels of mouth to work with — crop in before uploading.
  • Use flat, even lighting. Portraits lit softly from the front animate more cleanly than heavily stylised lighting with deep contrast across the face.
  • Avoid a mid-expression source. A photo caught mid-word or with the mouth already wide open is a strange starting point; a neutral or lightly smiling closed mouth is the most reliable base.
  • Match the background to the use. It comes along unchanged, so a plain wall suits corporate content and a styled set suits lifestyle content. That is decided when you pick the photo, not afterwards.

Recording the audio

  • Record in a quiet room, close to the microphone but not on top of it. Room echo and background noise blur the boundaries between sounds — exactly the information the model uses to drive the mouth.
  • Do not bake music under the voice. Add it afterwards in an editor: a voice under a music bed is harder to track, and mixing it in first means you cannot change the music without regenerating.
  • Watch the levels. Clipped, distorted audio syncs noticeably worse than a clean recording at a moderate level.
  • Speak at a normal pace. Very fast delivery compresses mouth shapes together; unnaturally slow delivery leaves the face idling between words.
  • Split long scripts. Two 20-second clips beat one 40-second clip — you can re-record a section without regenerating everything, and short segments hold attention better anyway.

What people actually use this for

Faceless channels with a consistent host

The hard part of an anonymous channel is that it has no face, and audiences attach to faces. A generated presenter solves that without putting you on camera. The workflow that holds up over dozens of uploads is to fix the presenter photo once — ideally generated with an AI image model, so the likeness belongs to no real person — and reuse that exact file every time. A slightly different photo each week reads as a slightly different person.

Spokesperson ads and landing-page video

Because the price is flat per video rather than per second, four script variants cost four times one video regardless of how long each runs. That changes how you approach ad creative: instead of committing to a single hook and buying media against it, generate four versions of the same pitch with different opening lines, run them all, and find out which one holds attention before spending real budget.

Multilingual redubs

One photo can front the same message in as many languages as you have recordings for. Generate one video per language track instead of subtitling a single video, and the mouth movement matches each language rather than contradicting it. The real constraint is translation quality, not the model: get the script translated properly and read by someone fluent, and the video follows. Timing shifts between languages, so each version will be a different length.

Training and internal video

Onboarding material, policy updates and internal announcements are the content nobody wants to book a studio for and everybody has to redo when a detail changes. A talking avatar makes the update cheap: change the paragraph, re-record that section, regenerate that clip, leave the rest untouched. Splitting the script by topic at the start is what makes that possible later — one long video means one long re-record.

The free alternative: PRUNA V

The dedicated Talking Avatar model is a premium feature, but it is not the only way to lip-sync a photo in VIBE. PRUNA V is a free model that also accepts an uploaded audio file and can produce talking-avatar style lip sync from a single photo. The differences are worth understanding before you pick one.

Talking AvatarPRUNA V
TierPremiumFree
InputPhoto + voice recording, nothing elseText prompt, or a photo, plus an optional audio file
Text promptNone — the model has no prompt fieldYes — you can describe the scene or the motion
Resolution480p or 720p720p
LengthMatches your recording, with no upper picker2 to 10 seconds
Aspect ratioInherited from the photo7 options including 16:9, 9:16 and 1:1
PricedFlat per video (720p costs more than 480p)Per second

The practical split: PRUNA V for short spoken clips of up to ten seconds, for testing whether a photo animates well before paying anything, and for cases where you also want a prompt to influence the scene. The Talking Avatar model when the script runs longer than ten seconds, when the output should preserve your photo’s framing exactly, and when a flat price per video makes testing several script variants worth it.

A third option exists if your goal is timing rather than speech. WAN 2.7 accepts an uploaded voice or music track — WAV or MP3, 3 to 30 seconds — and syncs the generated video to it, at 720p and up to 15 seconds. Reach for it when you want a clip cut to music rather than a face reciting a script.

Only animate a photo you have the right to use

Making a photograph of a real person appear to say words they never said is deceptive by default, and it does not stop being deceptive because the tool made it easy. Use your own face, a face whose owner has explicitly agreed to this specific use, or a character you generated. Do not animate a public figure, a competitor, a stranger from the internet, or a photo you found in a search result. The same applies to the voice: cloning someone’s voice without their permission carries the same problem as using their face, and in many jurisdictions the same legal exposure.

Consent, disclosure and where the line sits

This technology sits close to a genuine harm, so it is worth being specific rather than gesturing at "use responsibly".

Get permission in writing, for the actual use. A colleague who agreed to a photo for a team page has not agreed to front a paid ad campaign. If someone else’s likeness appears in commercial work, write down what they consented to, covering the specific channels and the specific message.

Never impersonate. Do not generate a video of a real person endorsing a product, stating an opinion, delivering news, or saying anything they did not say — including public figures, executives, and people who are no longer alive. This is the case that damages people, and it is why platforms and regulators are increasingly strict about synthetic likeness.

Disclose when it matters. Several platforms now require synthetic media to be labelled, and advertising rules in a number of markets require disclosure when a presenter is AI-generated. Labelling costs nothing and removes the entire category of complaint where an audience feels misled afterwards.

The safest habit is also the most convenient one: generate your presenter. An AI-generated portrait belongs to no living person, cannot withdraw consent, and can be reused indefinitely. For most commercial use cases it is both the lower-risk and the lower-friction option.

Common mistakes

  • Generating a five-minute video in one pass. Long single clips are harder to edit, more expensive to redo, and hold attention worse. Split by topic.
  • Choosing the photo before writing the script. The framing and background have to suit the message, and you only know the message once the script exists.
  • Skipping the 480p draft. One cheap generation tells you whether the photo animates well at all.
  • Mixing music into the voice track before generating. Add it afterwards, so you can change it without regenerating.
  • Expecting a full performance. You get convincing speech and natural head motion — not gestures, walking, or hand movement. Cut to other footage if the video needs more than a person talking.

AI talking avatars — frequently asked questions

Can I make a photo talk with AI?

Yes. The Talking Avatar model in VIBE takes a photo and an audio recording and generates a video of that face speaking, with mouth movement matched to the sound. There is no text prompt and no duration setting — the recording determines both what is said and how long the video runs. It works on iPhone and Android.

Is there a free AI talking avatar?

PRUNA V is a free model in VIBE that accepts an uploaded audio file and can produce talking-avatar style lip sync from a single photo. It generates clips of 2 to 10 seconds at 720p, so it suits short spoken lines. For longer scripts, the dedicated premium Talking Avatar model matches the video length to your recording with no ten-second limit.

Do I need a script or a prompt?

You need a script only in the sense that you need something to say into the recording — there is no text field in the app for the Talking Avatar model. Whatever is in the audio file is what the avatar says. If you would rather not use your own voice, generate or record the audio elsewhere and upload the file.

Can I use my own voice?

Yes, and it is the most common way people use it. Record on your phone in a quiet room, save as MP3, WAV or FLAC, and upload it. Your own voice avoids the consent questions that come with cloning someone else’s, and it usually sounds more natural than a synthetic track.

Is it legal to make someone’s photo talk?

It depends entirely on whose photo it is and what you make them say. Animating your own photo, or one you have explicit permission to use for that purpose, is fine. Generating a video of a real person saying something they never said — particularly a public figure, or anyone appearing to endorse a product — is deceptive, is prohibited by most platforms, and carries real legal exposure for likeness and publicity rights in many jurisdictions. Use your own face, a face you have written permission for, or an AI-generated character.

How long can a talking avatar video be?

The video is exactly as long as the audio you upload, since there is no duration picker. In practice, splitting a long script into segments of roughly 15 to 30 seconds is easier to work with: you can re-record one section without regenerating the rest, and shorter clips perform better on social platforms.

Should I generate at 480p or 720p?

Generate a 480p draft first to check the lip sync and the head movement, because it costs less. Regenerate at 720p for anything that will be seen on a larger screen — a product page, an ad, or a client deliverable. For a phone-sized social feed, 480p is often enough.

Does it work with illustrations and AI-generated faces?

Yes. The model animates illustrated characters and AI-generated portraits as well as photographs. A common workflow is to generate a spokesperson with one of the AI image models in VIBE, then use that image as the avatar — which also means the likeness belongs to no real person, removing the consent question entirely.

Make your first talking avatar

Download VIBE free on iOS and Android. Upload a photo and a voice recording to generate a lip-synced talking video — plus 35+ other AI video models, with free models included.

Download VIBE AI Video Generator on the App StoreGet VIBE AI Video Generator on Google Play