Guide
Image to Video AI: How to Turn a Photo Into a Video
Published By DreamClip
Image to video AI takes one still and renders the seconds that follow it. You upload a photo as the start frame, write one line about what moves, and the model keeps the face, the clothes and the room while it animates them. Open Kling V3 Turbo first for people and pets: the speed-first Kling, silent, 3 to 15 seconds, wide, tall or square. Open Veo 3.1 when the clip needs sound and light that follows the photo, such as a product still coming alive for an ad. Open Seedance 2.5 when the clip has to land on a second photo.
Every video model on DreamClip takes a start frame, so the choice is about length, sound, shape and price. The image to video page shows six real clips beside their photos; this guide is the method behind them.
One rule covers most of it: describe the change, not the picture. The model already has the picture. What it needs from you is the motion, the camera and, on the models that render it, the sound.
What image to video AI does, and what it does not
A text to video model invents the first frame and everything after it. An image to video model is handed the first frame and invents only what follows, which is why the result looks like your photo: the person, the room and the light are yours, and the model guesses how that scene moves next. With no photo yet, start from text to video instead.
It does not restore the photo: a blurry scan comes back as a blurry clip, so sharpen first, as our animate old photos guide does. It does not know what lies outside the frame, so a camera move that reveals the room invents the room. And it holds one motion well and three badly.
The photo is the first frame, not a loose reference. Crop it to the clip's shape before you upload.
Picking the start frame
The photo decides more than the prompt does. A sharp, well lit still with one clear subject animates well on every model; a dark, busy, low resolution one fails on all of them.
| The photo | What tends to happen | Do this instead |
|---|---|---|
| One subject, sharp, mid frame | What tends to happenClean motion, face and clothes held | Do this insteadUse it as it is |
| A crowd of ten | What tends to happenFaces at the back melt or swap | Do this insteadCrop to the two or three people who matter |
| A hand half out of frame | What tends to happenThe hand grows, or a sixth finger appears | Do this insteadCrop the hand out, or keep it still in the prompt |
| A dark, grainy phone shot | What tends to happenBlur and crawling noise | Do this insteadBrighten and enhance the still first |
| A square print for a vertical clip | What tends to happenA stretched or cropped subject | Do this insteadCrop to the clip's shape, and pick a model that renders it |
Match the clip's shape to the photo: a tall photo wants 9:16, a wide one 16:9. Most models here also render 1:1, which saves a crop on a square photo. Veo 3.1 and Gemini Omni Flash 1.1 render 16:9 and 9:16 only, and Minimax Hailuo 2.3 has no shape setting, so crop the photo to the shape you want before you upload it there.
Writing a motion prompt: describe the change, not the picture
A text to video prompt describes a scene. An image to video prompt describes a change to a scene the model can already see. Writing the picture again wastes the words and sometimes fights the photo. Leave out what is already in the frame, leave out negatives, and leave out the second scene: one clip is one moment.
The order to write a motion prompt in
- The subject's motion, first: she leans forward and blows out the candles.
- What else moves: the flames flicker and go out, smoke curls up.
- The camera, if it moves at all: handheld, a slow push in, or hold still.
- The sound, on a model that renders it: a room full of people clapping.
- The look, in two words at most: phone footage, cinematic.
The six clips on our image to video page were written this way, three on Kling V3 Turbo and three on Veo 3.1; their prompts are in the table further down. Ask for nothing sudden at the start: a jump or a fast turn in the opening frames tears the photo apart before the motion settles, and it is the commonest reason a clip ends on a different person.
The models that take a start frame on DreamClip
These ten cover the common jobs. Each cell is read from our model catalog: the lengths, sizes and sound options the studio offers today, not a maker's advertised maximum.
| Model | Maker | Clip lengths | Sizes | Sound | Start frame | Reference photos |
|---|---|---|---|---|---|---|
| Kling V3 Turbo | MakerKling | Clip lengths3 to 15 s | Sizes720p, 1080p | SoundNo | Start frameYes | Reference photosNo |
| Kling 3.0 | MakerKling | Clip lengths3 to 15 s | Sizes720p, 1080p | SoundYes, on by default | Start frameYes | Reference photosNo |
| Veo 3.1 (with audio) | MakerGoogle | Clip lengths4, 6 or 8 s | Sizes720p, 1080p | SoundYes, on by default | Start frameYes | Reference photosNo |
| Seedance 2.5 | MakerByteDance | Clip lengths4 to 30 s | Sizes480p, 720p, 1080p | SoundYes, on by default | Start frameStart and end | Reference photosNo |
| Seedance 2.0 Fast | MakerByteDance | Clip lengths4 to 15 s | Sizes480p, 720p | SoundYes, on by default | Start frameStart and end | Reference photosYes |
| Wan 3.0 | MakerAlibaba | Clip lengths2 to 30 s | Sizes480p, 720p, 1080p | SoundYes, on by default | Start frameStart and end | Reference photosNo |
| Gemini Omni Flash 1.1 | MakerGoogle | Clip lengths3 to 10 s | Sizes360p, 720p, 1080p, 4K | SoundAlways on | Start frameStart and end | Reference photosNo |
| Minimax Hailuo 2.3 | MakerMiniMax | Clip lengths6 or 10 s | Sizes768p | SoundNo | Start frameYes | Reference photosNo |
| Luma Ray v3.2 | MakerLuma | Clip lengths5 s | Sizes540p, 720p, 1080p | SoundNo | Start frameYes | Reference photosNo |
| Pika 2.2 | MakerPika | Clip lengths5 or 10 s | Sizes720p, 1080p | SoundNo | Start frameYes | Reference photosNo |
- KlingKling V3 Turbo3 to 15 s · 720p, 1080p
- KlingKling 3.03 to 15 s · 720p, 1080p
- GoogleVeo 3.1 (with audio)4, 6 or 8 s · 720p, 1080p
- ByteDanceSeedance 2.54 to 30 s · 480p, 720p, 1080p
- ByteDanceSeedance 2.0 Fast4 to 15 s · 480p, 720p
- AlibabaWan 3.02 to 30 s · 480p, 720p, 1080p
Start frame is universal here; end frame and reference photos are not. When the clip has to end on a second photo, open Seedance 2.5, Wan 3.0 or Gemini Omni Flash 1.1. For a subject that should appear in a scene the photo does not show, Seedance 2.0 Fast accepts as many as nine reference photos, cited in the prompt as @Image1, @Image2 and so on.
Which model for which photo: products, people, landscapes
The models differ less in whether they can animate a photo than in what they protect while they do it.
| The photo | Open first | Why |
|---|---|---|
| A person's face, a family moment | Open firstKling V3 Turbo | WhyHolds a face well for the first eight to ten seconds, silent, and renders square prints at 1:1 |
| A product on a surface | Open firstVeo 3.1 | WhyKeeps the photo's light and adds the sound of the thing: steam, a pour, a click |
| A pet mid action | Open firstKling V3 Turbo | WhyFast natural motion, four legs that land; Seedance 2.0 Fast is the second try |
| A landscape or a cityscape | Open firstWan 3.0 | WhyUp to 30 seconds for a slow drift, sound on by default; Gemini Omni Flash 1.1 for 4K |
| A car, a wedding, a dance | Open firstVeo 3.1 | WhyPhysics, reflections and two faces that hold, with the room's sound |
| A short loop for a feed | Open firstPika 2.2 or Luma Ray v3.2 | WhyShort silent clips, among the lowest cost per second here at 720p |
| A clip that must end on a second photo | Open firstSeedance 2.5 | WhyTakes a start and an end frame; so does Wan 3.0 |
The Kling page collects prompts for people and pets, the Veo page prompts that lean on sound and light, and the Seedance page the long takes and end frames.
Length, resolution and sound
Length is where the models part ways, and the model table further up lists every option. Veo 3.1 stops at 8 seconds, the Kling models at 15, and Seedance 2.5 and Wan 3.0 go to 30. A photo usually wants the short end: six seconds of a smile is a moment, fifteen is a loop.
Draft at 720p, or at 480p where it is offered, and render 1080p once the motion is right. On most models the bigger size costs more per second; Veo 3.1 prices 720p and 1080p the same. For 4K from a photo, Gemini Omni Flash 1.1 renders it natively, and so do Kling 3.0 4K and Kling O3 Omni 4K, both covered in our Veo 3.1 vs Kling 3.0 comparison.
Sound follows the model. Veo 3.1, Kling 3.0, Seedance 2.5, Seedance 2.0 Fast and Wan 3.0 render sound and have it on by default, with a switch to turn it off; on Veo 3.1 and Kling 3.0 a silent clip also costs less. Gemini Omni Flash 1.1 renders sound on every clip, and the model table marks the silent ones. For a photo of a person speaking, the better route is a talking avatar, which takes the portrait and an audio track and matches the mouth to the words.
Make the still first: generate, then animate
The best start frame is often one you made. A product shot that does not exist yet, a character you designed, a location you cannot visit: generate the still, judge it, then animate it. A still costs a fraction of a clip, so you pay for motion only once the frame is right.
- Open Nano Banana Pro and write the frame as a photograph: subject, setting, light, lens, and the clip's shape.
- Render two or three, and fix small things on the keeper with a photo edit rather than a re-roll: the image to image route takes the still and one line of change, and the Nano Banana page has frame prompts for it.
- Open the still as the start frame in the Video studio and describe only what moves.
Ten examples: six before-and-after pairs and four more to try
The first six are the real prompts behind the before-and-after pairs on our image to video page, which shows each photo beside the clip it became. The last four are new, written to the same rules and not rendered by us, so treat them as starting points. Use your own photo, keep the prompt, and change one thing at a time.
| The photo | The prompt | Model |
|---|---|---|
| A birthday cake, candles lit | The promptShe leans forward and blows out the candles, the flames flicker and go out, smoke curls up, she laughs and claps, the hands around her clap too, handheld phone footage | ModelKling V3 Turbo, the page's clip |
| A dog on a beach | The promptThe dog drops the ball, crouches, then bolts along the beach through shallow water, sand and spray kicking up, the camera panning to follow | ModelKling V3 Turbo, the page's clip |
| A coffee on a table | The promptSlow push in as steam rises and curls from the coffee, sunlight shifting across the table, a hand enters and lifts the mug | ModelVeo 3.1, the page's clip |
| A couple at a wedding | The promptThey sway and turn slowly, her dress swinging, she laughs and leans her head on his shoulder, the string lights twinkling, guests clapping softly at the edges, handheld | ModelVeo 3.1, the page's clip |
| A parked car | The promptThe car pulls away along the coast road, palms passing, the sun flaring off the chrome, camera tracking from behind | ModelVeo 3.1, the page's clip |
| Kids by a sprinkler | The promptThey land and bounce again, hair flying, the sprinkler sweeping across the frame, one of them shrieking with laughter, handheld | ModelKling V3 Turbo, the page's clip |
| A grandfather reading by a window | The promptHe looks up from the newspaper, takes off his glasses and smiles, the curtain moving in a breeze, the camera holds still | ModelKling V3 Turbo, a new prompt |
| A perfume bottle on marble | The promptA slow push in, light glints across the glass, a thin ribbon of mist drifts past, the bottle itself stays still | ModelVeo 3.1, a new prompt |
| A mountain lake at dawn | The promptMist drifts across the water, the reflection ripples once, the clouds move slowly left, the camera holds still | ModelWan 3.0, a new prompt |
| A city street at night | The promptRain starts to fall, the neon signs flicker, the reflections on the wet road shimmer, a slow push in | ModelGemini Omni Flash 1.1, a new prompt |
How to turn a photo into a video on DreamClip
- Open the Video studio with the start frame mode on: Kling V3 Turbo for people and pets, Veo 3.1 for products and sound, Seedance 2.5 when you need an end frame.
- Drop the photo on the start frame, cropped to the clip's shape.
- Write the motion line, one subject and one change, and pick the length and size.
- Run it. Credits are held when you press and settled only when a finished clip is delivered; an animation that fails is never charged.
- Judge the first second and the face. If either slips, calm the opening of the prompt or try the next model.
To choose between two models for one photo, run the same photo and motion line on each in the Video studio at its shortest length and smallest size, then render the keeper. Compare runs one text prompt on up to four video models side by side, each at its shortest length and smallest size in 16:9, with no charge for a failed lane; it takes no photo, so it picks a model for a motion, not for a picture.
Where image to video falls short
- Hands and teeth. A hand that waves and a mouth that opens wide are where every model still slips. Keep the hands still, and use a talking avatar for speech.
- A face turning past profile. The model has never seen the back of this head; it invents one. Keep turns under a quarter turn.
- A second person entering. Anyone not in the photo is a stranger the model makes up. If they matter, put them in the still first.
Luma's Dream Machine, Pika's own app, Pollo and Veed all offer a photo to video step, as of October 2026, each inside its own editor or subscription. DreamClip runs Luma Ray v3.2 and Pika 2.2 beside Kling, Veo and Seedance in one studio, so the comparison is a model switch rather than a new account.
Questions
Is image to video AI free on DreamClip?
Trying it costs nothing up front. Signing up needs no card, and a free account gets credits every day. Image to video is priced by the model, the length and the size, so a three second clip at the smallest size on a lower cost model fits a day's credits far more easily than an eight second Veo 3.1 clip. A clip that fails is never charged.
Which image to video AI keeps faces the same?
Kling V3 Turbo and Kling 3.0 hold a face better than most for the first eight to ten seconds, and Veo 3.1 keeps the photo's light and colour over its 4 to 8 seconds. Every model drifts past ten seconds from one photo, so keep clips of people short.
Can I use my photo as the first frame on Veo 3.1?
Yes. Open Veo 3.1 in the Video studio with the start frame mode on, upload the photo and describe the motion. Veo 3.1 renders 4, 6 or 8 seconds at 720p or 1080p in 16:9 or 9:16, with sound on by default.
Can I set the last frame as well as the first?
On some models. Seedance 2.5, Seedance 2.0, Seedance 2.0 Fast, Wan 3.0, Wan 2.7, Gemini Omni Flash 1.1, MiniMax H3, MiniMax H3 Max and PixVerse V6 take a start and an end frame on DreamClip. Kling V3 Turbo, Kling 3.0 and Veo 3.1 take a start frame only.
Can I use Sora for image to video here?
No. DreamClip does not run Sora, and OpenAI retired the Sora API on September 24, 2026. For a photo to video clip, the models here are Kling V3 Turbo and Kling 3.0, Veo 3.1, Seedance 2.5 and the others in the table above.
Does Kling 4.0 do image to video on DreamClip?
No. Kling's own site announces Kling 4.0 (as of October 2026), and DreamClip does not run it. Every Kling model sold here, from Kling V3 Turbo to Kling O3 Omni 4K, takes a start frame and none takes an end frame; if the clip must end on a second photo, open Seedance 2.5 or Wan 3.0.
Related
- GuideHow to Use Kling AI: Kling 3.0 for Text and Image to VideoHow to use Kling AI on Kling 3.0: text to video, image to video from a product photo, Motion Transfer, O3 Omni, 4K and the prompts that work.
- ComparisonVeo 3.1 vs Kling 3.0: Which AI Video Model to Pick in 2026Veo 3.1 vs Kling 3.0 on sound, clip length, 4K, start frames and motion, with the pick for ads, social clips and cinematic shots, and how to test both.
- GuideHow to Animate Photos With AI, Beyond Deep NostalgiaEnhance the scan, animate it from a start frame, then add a voice: a respectful method for old family photos, and what Deep Nostalgia does and does not do.
Sources
- Kling AI, the official site, checked