Upload a still image to an image-to-video model and describe a small, specific motion ("slow blink, slight smile, gentle push-in"). The model generates a 5 to 15 second clip that starts from your photo. That covers most jobs. For a person who speaks, use a talking photo (lip sync) tool with an audio track. To copy a dance or gesture from a real video, use motion control. To move between two specific images, use first and last frame.
Whatever the method, ask for less motion than you think you need. Every extra movement makes the model invent pixels it never saw, and that is where results turn uncanny.
Which photo animation method should you use?
| What you want | Method | Oakgen tool | Typical length |
|---|---|---|---|
| A still photo that "breathes" (blink, hair, light) | Image-to-video, subtle motion | AI video generator | 5 s |
| A slow cinematic camera move on a landscape or product | Image-to-video, camera move | AI video generator | 5 to 8 s |
| A person or pet doing a new action | Image-to-video, action prompt | AI video generator | 5 to 10 s |
| A person saying specific words | Talking photo (lip sync) | Talking Photo | Length of your audio, up to 59 s on VEED Fabric |
| A person copying a dance or gesture from a video | Motion control | Kling V3 Pro Motion Control | Up to 30 s |
| A transition from one image to another (before/after, young to old) | First and last frame | Kling models with an End Frame field, or Seedance 2.0 | 4 to 10 s |
Old family photos and portraits want the first row. Most bad animated photos come from people picking row three when they wanted row one.
Model names, durations, input fields, and credit costs are taken from the settings Oakgen had in place in September 2026, at 1 USD = 260 credits. We have not benchmarked output quality between models.
Method 1: Image-to-video (the default way to animate a photo)
Image-to-video models take your photo as the first frame and generate what happens next. On Oakgen the main options for this are Kling V3 Pro (3 to 15 seconds, optional generated audio), Seedance 2.0 (4 to 15 seconds, 480p to 1080p), and Kling O1 (3 to 10 seconds).
There are really three jobs hiding inside "animate this photo," and each needs a different prompt.
Camera move
The subject stays still. The camera moves. This is the safest animation there is, because the model barely has to invent anything. It works for landscapes, architecture, products, and group photos where you do not want faces to change.
Formula: [camera move] + [speed] + [what stays still] + [one ambient detail]
Slow push-in toward the house, camera moves forward about one meter,
the scene itself stays still, light clouds drift across the sky.
Good camera verbs: slow push-in, slow pull-back, gentle pan left, slight orbit, handheld drift, rack focus. One per clip.
Subtle motion
The camera is locked and the subject does something tiny. This is the right call for portraits, old photos, and pets. Deep Nostalgia, the MyHeritage feature that launched in February 2021 and quickly had people posting animated ancestors on Twitter, worked by applying short "driver" performances of smiling, blinking, and turning the head. That restraint is a large part of why its results felt real to people.
Formula: [locked camera] + [one or two micro-actions] + [one environmental motion] + [keep identity]
Static camera. She blinks slowly and her smile widens slightly.
A few strands of hair move in a light breeze. Her face, clothing,
and the background stay exactly the same.
Action
The subject does something new: a dog runs toward the camera, a person turns and walks away, a product lid opens. Image-to-video is most useful here, and also fails most often here. The further the action takes the subject from the original pose, the more the model has to guess about hands, backs of heads, and hidden sides.
Formula: [subject] + [one action with a clear start and end] + [camera behavior] + [pace]
The golden retriever stands up, shakes its fur, then trots toward the camera.
Camera stays at dog eye level and slowly backs away. Natural, real-time pace.
Image-to-video is the most flexible of the four methods. It works on any subject from a single input and needs no audio. What you give up is control over exact timing, identity holds poorly under large motion, and you cannot make someone say specific words.
What image-to-video costs on Oakgen
| Model | Setting | Credits (as of September 2026) |
|---|---|---|
| Kling O1 | 5 s | 146 |
| Seedance 2.0 | 5 s, 480p | 175 |
| Seedance 2.0 | 5 s, 720p | 394 |
| Kling V3 Pro | 5 s | 437 |
Test at the cheapest setting that shows the motion, then rerun the keeper at a higher tier. A photo animation usually needs two or three attempts, so budget for several runs.
Animate Your First Photo
Upload a photo, pick Kling or Seedance, and generate a 5-second test clip.
Method 2: Talking photo (lip sync)
A talking photo takes one face image and one audio track and animates the mouth, jaw, head, and expression to match the speech. It is a different kind of model from image-to-video: the audio drives the motion, so timing is exact.
Oakgen's Talking Photo tool currently offers three lip-sync models:
| Model | Inputs | Resolution | Notes | About this many credits for 10 s of audio |
|---|---|---|---|---|
| Hedra Lipsync | Face image + audio | 480p or 720p, 9:16, 16:9 or 1:1 | Lowest per-second rate | 92 |
| InfiniteTalk | Face image + audio | 480p or 720p | Suited to longer messages | 156 |
| VEED Fabric 1.0 | Face image + audio | 480p or 720p | Audio capped at 59 s | 208 |
You can upload or record audio (the tool accepts up to 2 minutes), or type a script of up to 2,000 characters and generate the voice inside the tool. The face image should be a clear, front-facing photo under 10 MB. Our step-by-step Talking Photo walkthrough covers the interface.
What makes a talking photo work:
- A face at least a third of the frame height, looking roughly at the camera.
- A closed or neutral mouth in the source photo. An open-mouth laugh gives the model conflicting information.
- Clean audio with no music underneath. Background music confuses the mouth timing.
- Short sentences. A 20-second message looks more natural than a 60-second speech, because head motion starts to loop on longer clips.
You get exact speech and predictable timing, and it works on old photos and illustrations. On the downside, the body barely moves, output tops out at 720p, and it raises the biggest consent questions of the four methods (more on that below).
Method 3: Motion control from a reference video
Motion control takes your photo plus a short reference video of someone moving, then makes the person in the photo perform that exact motion. It is how most AI dance videos are made, because a text prompt cannot describe choreography precisely enough.
On Oakgen this runs on Kling V3 Pro Motion Control and Kling 2.6 Pro Motion Control. Kling V3 Pro has two modes: match the video's orientation (reference clips up to 30 seconds) or match the image's orientation (up to 10 seconds). Cost follows the length of the reference video, about 437 credits for a 10-second reference as of September 2026.
The Kling motion control guide covers settings and prompts in depth, and the AI dance video guide walks through the full photo-to-dance workflow.
Best for: precise, repeatable movement. It is the only method that handles full-body choreography well. Watch for: you need a clean reference video, the photo and the video should have a similar framing (full body with full body), and fast spins still smear faces.
Method 4: First and last frame interpolation
Here you give the model two images, a start and an end, and it generates the motion in between. This is the right method when you know exactly where the clip should land: a before and after renovation, a sketch becoming the finished painting, a baby photo morphing into an adult portrait, a product closed then open.
Several Oakgen video models accept an optional End Frame, including Kling V3 Pro, Kling 2.6 Pro, and Kling O1. Seedance 2.0 takes a last-frame image too.
Prompt formula: describe the transition itself. The model already sees both images.
Smooth continuous transition. The sketch lines fill with color from left
to right until the finished painting is revealed. Camera stays locked.
Keep the two frames close in composition. If the start frame is a head-and-shoulders portrait and the end frame is a full-body shot in a different room, the model has to invent a camera move and a scene change at once, and the middle of the clip turns to mush. Our first frame vs last frame guide explains when a last frame helps and when references work better.
Because you control the ending, this method suits reveals and transitions. It needs two good images, and large differences between them produce morphing artifacts.
Prompt formulas by photo type
Copy these, then change the details in brackets. Each one is written for image-to-video unless noted.
Portrait (modern photo)
Static camera, shallow depth of field. [Name/the woman] blinks naturally,
takes a slow breath, and gives a small closed-mouth smile. Slight head tilt
to her left. Background softly out of focus and still. Keep her face identical.
Old family photo
Locked-off camera. The man in the photo blinks slowly and his expression
softens into a faint smile. Very slight head movement. No other motion.
Preserve the original film grain, black-and-white tone, and his exact features.
Keep old photos to 5 seconds. Longer clips give the face more time to drift away from the original.
Product shot
Slow 20-degree orbit around the [bottle] on the table. Soft studio light,
a gentle highlight slides across the glass. The product, label, and
text stay sharp and unchanged.
Text on labels is the first thing to break. If the label matters, use a camera move only and never ask the product itself to move.
Landscape
Slow push-in toward the mountains. Clouds drift left to right, the lake
surface ripples lightly, grass sways in a soft wind. Golden hour light
stays constant.
Pet
Camera at eye level, static. The cat blinks slowly, ears twitch once,
whiskers move as it breathes. It stays sitting in the same pose.
For action with pets, one verb only: "tilts its head," "wags its tail," "yawns." Pets rarely survive "runs, jumps, and catches the ball" in one clip.
Artwork or illustration
Keep the painted style and brush texture exactly. Gentle parallax:
foreground flowers sway, the woman's hair and dress move in a breeze,
the sky shifts slowly. Camera does not move.
Add "keep the painted style" or the model may slide toward photorealism halfway through.
Why less motion looks more real
The model only knows what is in your photo. A slight smile uses pixels it already has. A full head turn requires it to invent an ear, the back of a jaw, and hair it has never seen, and that invented detail is where your grandmother stops looking like your grandmother.
| Too much | Better |
|---|---|
| "She laughs, turns around, and waves at the camera" | "She smiles and her eyes crinkle slightly" |
| "The dog runs across the park and jumps" | "The dog's ears perk up and it tilts its head" |
| "Dramatic camera sweep around the whole family" | "Very slow push-in, family stays still" |
| "The car drives away down the road" | "Camera slowly pans along the car, reflections move across the paint" |
| "Grandpa stands up and walks to the window" | "Grandpa blinks and looks slightly toward the window" |
A good test: if a person could do the motion without their face leaving the frame or changing angle much, the model can usually do it too.
How to fix common artifacts
| Problem | Likely cause | Fix |
|---|---|---|
| Face changes into a different person | Too much head rotation, low-res source | Reduce motion, upscale the photo first, add "keep her face identical" |
| Melting or extra fingers | Hands move in the prompt | Keep hands out of the action or crop them out |
| Mouth moves when nobody should speak | Model added speech or generated audio | Add "mouth closed," turn off audio generation |
| Background warps or breathes | Model treats the whole frame as flexible | Add "background stays completely still," use a locked camera |
| Text or logo scrambles | Product itself moving | Camera move only, no object motion |
| Clip goes slow-motion or frozen | Prompt too vague | Name one clear action and "natural real-time pace" |
| Color shifts mid-clip | Old or heavily compressed photo | Restore and upscale before animating |
| Talking photo looks robotic | Long audio, music in the track | Shorter lines, clean voice-only audio |
When a clip fails twice with the same prompt, change the photo or the method, not the adjectives. Rewording "gently" to "softly" rarely fixes a structural problem.
Restore and upscale old photos before you animate them
Animation magnifies every flaw in the source. Scratches flicker, blur turns to smear, and a low-resolution face gives the model almost nothing to hold onto.
A reliable order for old photos:
- Scan at 600 DPI or higher, or photograph the print flat in even daylight.
- Restore the face with Oakgen's Image Restorer. It uses GFPGAN, a face restoration model, and costs about 1 credit per image as of September 2026. Compare the restored face against the original: restoration can smooth skin and subtly change features, and a small change gets amplified once the face moves.
- Upscale with the Image Upscaler if the face is still small in the frame.
- Animate with a subtle-motion prompt, 5 seconds.
- Upscale the video afterward with the Video Upscaler if you need a sharper final file.
Our guide to restoring and upscaling old family photos covers scanning and damage repair in detail.
The Image Restorer is face-focused, though. Heavy tears, missing corners, and damage across the background may still need manual retouching or an image editor pass before animation.
Restore the Photo First
Clean up faces on old, faded, or blurry prints before you animate them.
Consent and ethics: animating photos of real people
A moving face is far more persuasive than a still one.
When MyHeritage launched Deep Nostalgia in 2021, it deliberately left speech out "in order to prevent abuse, such as the creation of deepfake videos of living people," and acknowledged that some people find the feature magical "while others find it creepy," according to the BBC's coverage. Today's talking photo tools do add speech, which makes the questions sharper.
Rules we suggest:
- Living people: only animate someone who has agreed to it. Never make a real person appear to say something they did not say, and never publish it as if it were real footage.
- Deceased relatives: ask close family before sharing. Reactions vary a lot, and a grieving parent may feel very differently from a curious cousin. Keep motion subtle, and be careful with speech. A short message based on something they actually wrote or said is very different from inventing new words.
- Public figures: satire and commentary are one thing; content that could be mistaken for a real statement is another. Label it clearly.
- Children: do not animate photos of other people's children without a parent's permission.
- Labels: when you share an animated photo, say it is AI-animated. It costs nothing and prevents confusion.
If you are making a tribute video, our talking photo memorial video guide walks through scripting and voice choices with these concerns in mind.
How to animate a photo in six steps
- Decide the motion first, using the decision table above.
- Prep the photo: crop to the aspect ratio you want, restore and upscale if it is old or small.
- Write one sentence of motion using the formula for your photo type.
- Generate a 5-second test at a low tier.
- If the face drifts, cut the motion in half. If nothing moves, name one clearer action.
- Rerun the winner at a higher resolution, then upscale if needed.
Make a Photo Talk
Pair a face photo with your own voice, a recording, or a typed script and generate a lip-synced clip.
Sources
- MyHeritage Blog, "New: Introducing Deep Nostalgia -- Animate the Faces in Your Family Photos," February 25, 2021: https://blog.myheritage.com/2021/02/new-animate-the-faces-in-your-family-photos/ (accessed 2026-09-26)
- BBC News, "MyHeritage offers 'creepy' deepfake tool to reanimate dead," February 26, 2021: https://www.bbc.com/news/technology-56210053 (accessed 2026-09-26)
- Oakgen model configuration for video, lip-sync, and image restoration models, including durations, input fields, and provider costs converted at 1 USD = 260 credits (checked 2026-09-26)