An AI lip-sync clip succeeds when the voice, mouth movement, face, and edit feel like one performance. The reliable path is to separate two jobs: create the visual performance, then decide whether its speech should come from the generation itself or from a dedicated lip-sync pass.
Oakgen's current Seedance 2.0 text-to-video workflow accepts audio references, but its form does not expose a separate generated-audio switch. That makes audio useful for directing the scene, not a reason to promise exact mouth matching. For presenter ads and talking portraits, Oakgen also has UGC workflows with dedicated lip-sync models. Pick the route that matches the asset you already have.
| Starting point | Best route | Why |
|---|---|---|
| A script and character references | Generate the video first | You still need the performance, setting, and camera |
| An approved portrait and voice file | Talking-photo or UGC lip-sync workflow | The identity is fixed and speech is the main job |
| An approved video with the wrong dialogue | Dedicated lip-sync or dubbing pass | Preserve the edit and change only the speaking performance |
| A music track with a few sung close-ups | Hybrid workflow | Generate performance shots, then apply lip-sync only where the mouth is readable |
Seedance 2.0 accepts audio references in the current text-to-video workflow. Oakgen does not show a user-facing generated-audio switch for this model. Oakgen's UGC lip-sync modes accept an image or approved visual plus an audio file for a more specific talking-character job.
Choose generation or post-production lip sync
Use video generation when the scene still needs invention. The prompt controls posture, expression, action, environment, camera, and pacing. An audio reference may guide rhythm or performance, but the result still needs a full visual and sound review.
Use a dedicated lip-sync pass when the source picture already works. This route is easier to judge because fewer elements are changing. It also keeps a successful camera move, product interaction, or background performance intact while replacing the spoken line.
A useful rule: if you would be upset to lose the current picture, do not reroll the entire shot just to change its dialogue.
Prepare the audio before generating
Clean speech gives every lip-sync system a clearer target. You do not need a studio master, but you do need an intelligible line.
- Use one speaker at a time.
- Remove music and loud effects from the speech file when possible.
- Avoid clipping, harsh echo, and heavy noise reduction artifacts.
- Trim long silence before and after the line.
- Keep the spoken line short enough to fit one shot naturally.
- Export a common format accepted by the selected workflow and confirm it plays before uploading.
Do not optimize around an invented technical rule such as one mandatory sample rate or a guaranteed no-drift duration. Input requirements differ by model and provider. The current form is the source of truth.
If the voice does not exist yet, create and approve it in the AI voice generator before making video. Approval means checking pronunciation, pace, emphasis, tone, and the legal right to use the voice.
Workflow 1: a short lip-sync meme
A meme needs one readable reaction and one line. Extra camera movement usually weakens the joke.
- Write a line that fits the intended shot without rushing.
- Generate or select a front-facing character image with an unobstructed mouth.
- Record or generate the voice and export the isolated speech.
- Use a talking-photo or UGC lip-sync workflow when exact dialogue is the main requirement.
- Add captions in the editor, where spelling and timing remain under your control.
Use a visual prompt like this for the source portrait:
Chest-up portrait of a tired office worker in a plain break room, facing the camera, neutral head angle, mouth fully visible, soft overhead light, restrained deadpan expression, vertical composition. No text, microphone, hand over face, extreme profile, or exaggerated open mouth.
The joke should come from the line and expression. Do not ask the model to generate readable caption text inside the frame.
Workflow 2: a music-led video
Most music-video shots do not need visible lip sync. Reserve it for close and medium performance shots where viewers can read the mouth. Use non-speaking shots for movement, locations, product details, crowd reactions, and transitions.
Build the sequence in three lanes:
- Performance: the singer faces camera and delivers a short approved section.
- Story: character action advances the visual idea without readable speech.
- Texture: objects, environments, hands, light, and motion provide edit points.
For each visible singing shot, use the isolated vocal section as the lip-sync input. Put the final music mix over the completed edit afterward. This keeps drums and instruments from competing with the speech or vocal signal during the mouth-animation step.
Seedance 2.0 can help create reference-led performance and story shots. A dedicated lip-sync mode is the safer finishing choice when the vocal match is the acceptance test. Use the music generator for an original track when needed, then keep a clear record of which audio master belongs to each edit.
Workflow 3: a recurring presenter Reel
Consistency comes from a small production system, not one long prompt. Save an approved presenter pack containing:
- a front-facing identity image;
- one alternate angle;
- the exact wardrobe description;
- the room, light direction, and framing;
- an approved voice;
- the caption style and safe area;
- negative constraints for features that must not change.
Break a 30-second Reel into separate beats: hook, explanation, example, and next step. Generate or lip-sync each beat as a short shot. This gives you cleaner edit points and lets you replace one weak line without rebuilding the full Reel.
For a product-led presenter video, open Oakgen's UGC ad generator. Use the general AI video generator when the job depends more on scene creation, camera, or reference-directed action than on a talking avatar.
Give every input one job
Labeling prevents reference conflicts. A simple production note can look like this:
| Asset | Controls | Must not control |
|---|---|---|
| Identity image | Face, hair, age presentation | Camera path or background |
| Wardrobe image | Garment shape and color | Face identity |
| Voice file | Words, pace, vocal performance | Visual style |
| Environment image | Room, palette, light | Mouth shape |
| Motion reference | Gesture or camera rhythm | Product geometry |
When two inputs disagree about the same feature, remove one or state which asset wins. Adding more references does not repair a contradictory brief.
Review the result in four passes
1. Speech
Watch once with sound. Check whether starts, stops, pauses, and emphasis agree with the visible performance. If a word feels late, identify the exact time instead of asking for “better sync.”
2. Face
Mute the clip and watch the mouth, jaw, cheeks, eyes, and head together. A technically moving mouth can still look pasted on when the rest of the face stays frozen.
3. Identity and product
Compare the first, middle, and final frames with approved references. Reject face drift, changing wardrobe, altered packaging, extra fingers, or a product that changes shape.
4. Edit and delivery
Check caption room, crop safety, first-frame clarity, ending pose, and any platform disclosure. View the exported file on a phone before publishing.
| Symptom | Change one thing first | Do not do |
|---|---|---|
| Mouth timing feels late | Trim leading silence or shorten the line | Change the face, camera, and audio together |
| Face looks stiff | Ask for restrained blinks and head motion | Add several large gestures |
| Identity changes | Use one stronger identity anchor | Add many conflicting portraits |
| Words are hard to understand | Replace the voice file with cleaner speech | Cover the problem with louder music |
| Captions are wrong | Add captions after generation | Ask the video model to render final typography |
Rights, consent, and disclosure
Do not clone or animate a real person's face or voice without permission. Confirm music and footage rights before commercial use. Political, medical, financial, and news-like media need extra care because an altered speaking performance can mislead viewers even when the image quality is imperfect.
Keep source files, consent records, prompts, and final exports together. Add synthetic-media disclosure where a platform, client, or audience context calls for it.
A practical Oakgen route
Start with the asset you already trust:
- approved portrait plus voice: use the UGC ad generator;
- references plus a new scene: use the AI video generator;
- missing narration: create it with the AI voice generator;
- missing visual anchor: build it with the AI image generator.
Seedance 2.0 is one option for reference-led video, not a blanket promise of exact lip sync. When the mouth performance is the core requirement, compare it with a dedicated lip-sync model using the same image and audio, then keep the result that passes the review checklist.
Build a controlled lip-sync test
Bring one approved portrait and one clean voice file to Oakgen, then compare the workflow against a clear acceptance checklist.
