Tutorial

AI lip sync, explained: how to make characters speak.

Sculpted glass lips with a blue audio waveform passing through

AI lip sync is the step that makes a generated or animated character’s mouth match a spoken audio track. Two pipelines drive it. Audio-driven: a model reads the waveform and moves the mouth to fit. Performance-driven: an actor’s face on camera feeds mouth shape, timing, and micro-expression onto the character. Both end at the same deliverable, a shot where the words look spoken rather than pasted on. Which one you reach for depends on how much live performance the shot needs.

The short version
  • Two pipeline families: audio-driven (voice track first, model animates the mouth) and performance-driven (an actor’s face drives the character).
  • Lock the voice track before you sync. Re-generating the voice after sync means redoing the shot, not adjusting it.
  • One cloned or designed voice per character, saved to a voice library, keeps a project’s cast consistent from scene to scene.
  • Sync sits after shot generation and before the edit. Check for mouth-interior mush, sibilant timing, a frozen upper face, drift on long takes, and accents before you cut it into the timeline.

How does AI lip sync work?

Every AI lip sync tool takes two inputs, a face and a voice, and produces one output: a shot where the mouth appears to make the sounds on the track. The pipelines differ in where the performance comes from.

Audio-driven pipelines start with a finished voice track, generated or recorded, and feed it to a model trained to map audio to mouth shapes. The model reads phonemes and timing from the waveform and animates the lips, jaw, and often the surrounding face. This is the workflow behind avatar tools like HeyGen: you supply speech, the model does the rest, no camera needed.

Performance-driven pipelines start with an actor in front of a camera. Rather than infer mouth movement from sound alone, the model reads the actor’s face, its timing, blinks, and micro-expression, and transfers that performance onto the target character. Runway’s Act-One works this way: one camera feed drives the whole character performance, not the mouth alone. The tradeoff is direct. Audio-driven is faster and needs no actor. Performance-driven carries more of a human performance across, and needs a human to capture it.

AUDIO-DRIVEN Voice track Sync model Finished shot PERFORMANCE-DRIVEN Actor capture Performance transfer Finished shot Different inputs drive the mouth; both pipelines produce the same deliverable.
Audio-driven pipelines animate the mouth from a waveform. Performance-driven pipelines transfer an actor’s face. Either way, the output is a finished, synced shot.

What happens when the model generates the voice with the shot?

A third route skips the sync pass. Models with native audio produce the picture and the speech in one sample, so no waveform has to be matched to a mouth afterward. The tradeoff is control: when the timing misses, there is no sync pass to re-run, only a whole shot to generate again.

Here is one line, one pass, so you can judge the result against the checklist further down. The prompt names the shot, the lighting, the spoken line in quotes, and the beat after it.

Prompt · Veo 3.1 Fast, native audio
Medium close-up of a radio host in a small studio at night, warm tungsten practicals. She leans to the microphone and says: 'Signal over noise, every single week.' Then she smiles. Static camera.
Veo 3.1 Fast, native audio: the voice and the mouth came out of one generation pass, a third route beside audio-driven and performance-driven sync. Listen for the sibilants in “Signal” and “single” and watch whether they land on the shape that makes them. The brows and cheeks lift into the smile, so this is not a jaw moving under a frozen face. Two caveats: the three-quarter angle hides half the mouth, which is the forgiving case, and she drops her head on the last beat, so the line ends away from camera. Six seconds shows you sync quality. It does not replace a locked voice track, because every new pass returns a new voice.

Should you generate the voice before or after the video?

Before, always. Lock the voice track, then sync to it. A sync pass, in either pipeline family, is built against one specific waveform: its phoneme timing, its pauses, its pacing. Swap the voice after the model has animated the mouth to that take, and the mouth matches audio nobody will hear. That sends the shot back through sync and everything after it.

This is picture lock applied one department earlier. Nothing downstream should assume the upstream piece is still moving. Treat the voice track as locked the moment it goes to sync, the same boundary our guide to AI filmmaking draws around a locked shot. If the line reading is wrong, fix it in the voice step and re-run sync.

How do you keep a character’s voice consistent across a project?

Assign one cloned or designed voice per character, and keep it in a voice library rather than generating a fresh voice each time that character has a line. A voice library is a saved, reusable voice profile: every scene, every episode, every reshoot pulls the same voice. Voice tools built for production work, ElevenLabs among them, let you save a cloned or designed voice once and call it by name across a project.

Skip this step and the drift shows fast. A character’s tone shifts scene to scene, an accent wanders, a line reading in episode four does not sound like the person from episode one. For a recurring cast, treat the voice library the way you treat a model sheet for a character’s face: reference material the whole team draws from. Our glossary covers how voice cloning and voice design differ if you are choosing between them.

Where does lip sync belong in your pipeline?

After the shot is locked, before the edit. The order in practice: write the script, generate and lock the voice track, generate the shot, run the sync pass, hand the synced shot to the edit. Sync needs a locked voice and a finished shot to exist before it runs, and the edit needs sync finished, so its position in the pipeline is fixed.

SHOT PIPELINE Script Voice track Shot generation Sync pass Edit Sync starts only after the voice locks and the shot exists. The edit starts only after sync.
Sync is the hub of the pipeline: two locked inputs in, one synced shot out.

What failure modes should you check before you ship?

Lip sync fails in a short list of recognizable ways. Check for these on every synced shot before it goes into the edit, the way you check focus and exposure.

  • Mouth interior mush. The lips move, but the teeth and tongue behind them stay soft or smear together instead of forming the shapes a mouth makes on plosives and open vowels.
  • Sibilant timing. Sounds like s and sh land a frame or two off the mouth shape that should produce them. It reads as sloppy even when the rest of the take is clean.
  • Jaw-only movement. The jaw and lips animate while the eyes, brows, and cheeks stay still. Fastest way to make a character read as a puppet.
  • Sync drift on long takes. The mouth tracks the audio for the first few seconds and falls out of step as the clip runs on.
  • Accents and fast speech. Models trained mostly on one accent or a moderate speaking pace miss timing on quick dialogue and unfamiliar phoneme patterns.

Any one of these breaks a viewer’s attention on a close-up, even when the rest of the frame is clean. For more on what still needs a human hand in a generative pipeline, our post on the shift toward everything-models covers how much of production these systems can carry alone.

When should you cut around a lip sync shot?

When sync is not holding and you are out of time to fix it. Reaction shots and off-screen lines are cheaper than bad sync. If a line can play over a listener’s face, a cutaway, or an off-screen delivery, that edit choice costs nothing next to a shot whose mouth visibly disagrees with its audio. Save the on-mouth close-ups for lines where the audience needs to watch the character speak, and build coverage so you have an escape hatch when a sync pass does not clear the checklist above.

Questions creatives ask

How does AI lip sync work? It takes a face and a voice and returns a shot where the mouth matches the audio. Audio-driven models read phoneme timing from a waveform and animate the mouth to fit; performance-driven models capture an actor’s face on camera and transfer that performance, timing and all, onto the character.

Should you generate the voice before or after the video? Before. Lock the voice track, then run sync against it. Re-generate the voice afterwards and the animated mouth matches a take nobody will hear, so the shot goes back through sync.

How do you keep a character’s voice consistent across a project? Save one cloned or designed voice per character in a voice library and reuse it for every line that character has. That single saved profile keeps tone, accent, and cadence from drifting scene to scene.

Why does AI lip sync look uncanny? Usually one of five failures: mush where the mouth interior should show teeth and tongue, sibilants landing off-beat, a jaw that moves while the rest of the face stays frozen, sync drift over a long take, or a model missing timing on an accent or fast speech. Check for those directly and you catch most of what makes a synced shot read as off.

Learn this beside the people building it.

Membership is free. Masterclasses from industry leaders, hackathons where you finish something the same day, and mentor circles matched to what you want to learn. For engineers and creatives alike, across film, design, image, sound, and story.