Guide

Image to Video AI: How to Turn a Photo Into a Video

Published By DreamClip

Image to video AI takes one still and renders the seconds that follow it. You upload a photo as the start frame, write one line about what moves, and the model keeps the face, the clothes and the room while it animates them. Open Kling V3 Turbo first for people and pets: the speed-first Kling, silent, 3 to 15 seconds, wide, tall or square. Open Veo 3.1 when the clip needs sound and light that follows the photo, such as a product still coming alive for an ad. Open Seedance 2.5 when the clip has to land on a second photo.

Every video model on DreamClip takes a start frame, so the choice is about length, sound, shape and price. The image to video page shows six real clips beside their photos; this guide is the method behind them.

One rule covers most of it: describe the change, not the picture. The model already has the picture. What it needs from you is the motion, the camera and, on the models that render it, the sound.

What image to video AI does, and what it does not

A text to video model invents the first frame and everything after it. An image to video model is handed the first frame and invents only what follows, which is why the result looks like your photo: the person, the room and the light are yours, and the model guesses how that scene moves next. With no photo yet, start from text to video instead.

It does not restore the photo: a blurry scan comes back as a blurry clip, so sharpen first, as our animate old photos guide does. It does not know what lies outside the frame, so a camera move that reveals the room invents the room. And it holds one motion well and three badly.

The photo is the first frame, not a loose reference. Crop it to the clip's shape before you upload.

Picking the start frame

The photo decides more than the prompt does. A sharp, well lit still with one clear subject animates well on every model; a dark, busy, low resolution one fails on all of them.

The photoWhat tends to happenDo this instead
One subject, sharp, mid frameWhat tends to happenClean motion, face and clothes heldDo this insteadUse it as it is
A crowd of tenWhat tends to happenFaces at the back melt or swapDo this insteadCrop to the two or three people who matter
A hand half out of frameWhat tends to happenThe hand grows, or a sixth finger appearsDo this insteadCrop the hand out, or keep it still in the prompt
A dark, grainy phone shotWhat tends to happenBlur and crawling noiseDo this insteadBrighten and enhance the still first
A square print for a vertical clipWhat tends to happenA stretched or cropped subjectDo this insteadCrop to the clip's shape, and pick a model that renders it

Match the clip's shape to the photo: a tall photo wants 9:16, a wide one 16:9. Most models here also render 1:1, which saves a crop on a square photo. Veo 3.1 and Gemini Omni Flash 1.1 render 16:9 and 9:16 only, and Minimax Hailuo 2.3 has no shape setting, so crop the photo to the shape you want before you upload it there.

Writing a motion prompt: describe the change, not the picture

A text to video prompt describes a scene. An image to video prompt describes a change to a scene the model can already see. Writing the picture again wastes the words and sometimes fights the photo. Leave out what is already in the frame, leave out negatives, and leave out the second scene: one clip is one moment.

The order to write a motion prompt in

  1. The subject's motion, first: she leans forward and blows out the candles.
  2. What else moves: the flames flicker and go out, smoke curls up.
  3. The camera, if it moves at all: handheld, a slow push in, or hold still.
  4. The sound, on a model that renders it: a room full of people clapping.
  5. The look, in two words at most: phone footage, cinematic.

The six clips on our image to video page were written this way, three on Kling V3 Turbo and three on Veo 3.1; their prompts are in the table further down. Ask for nothing sudden at the start: a jump or a fast turn in the opening frames tears the photo apart before the motion settles, and it is the commonest reason a clip ends on a different person.

The models that take a start frame on DreamClip

These ten cover the common jobs. Each cell is read from our model catalog: the lengths, sizes and sound options the studio offers today, not a maker's advertised maximum.

Read from the DreamClip catalog when this page was built.
ModelMakerClip lengthsSizesSoundStart frameReference photos
Kling V3 TurboMakerKlingClip lengths3 to 15 sSizes720p, 1080pSoundNoStart frameYesReference photosNo
Kling 3.0MakerKlingClip lengths3 to 15 sSizes720p, 1080pSoundYes, on by defaultStart frameYesReference photosNo
Veo 3.1 (with audio)MakerGoogleClip lengths4, 6 or 8 sSizes720p, 1080pSoundYes, on by defaultStart frameYesReference photosNo
Seedance 2.5MakerByteDanceClip lengths4 to 30 sSizes480p, 720p, 1080pSoundYes, on by defaultStart frameStart and endReference photosNo
Seedance 2.0 FastMakerByteDanceClip lengths4 to 15 sSizes480p, 720pSoundYes, on by defaultStart frameStart and endReference photosYes
Wan 3.0MakerAlibabaClip lengths2 to 30 sSizes480p, 720p, 1080pSoundYes, on by defaultStart frameStart and endReference photosNo
Gemini Omni Flash 1.1MakerGoogleClip lengths3 to 10 sSizes360p, 720p, 1080p, 4KSoundAlways onStart frameStart and endReference photosNo
Minimax Hailuo 2.3MakerMiniMaxClip lengths6 or 10 sSizes768pSoundNoStart frameYesReference photosNo
Luma Ray v3.2MakerLumaClip lengths5 sSizes540p, 720p, 1080pSoundNoStart frameYesReference photosNo
Pika 2.2MakerPikaClip lengths5 or 10 sSizes720p, 1080pSoundNoStart frameYesReference photosNo

Start frame is universal here; end frame and reference photos are not. When the clip has to end on a second photo, open Seedance 2.5, Wan 3.0 or Gemini Omni Flash 1.1. For a subject that should appear in a scene the photo does not show, Seedance 2.0 Fast accepts as many as nine reference photos, cited in the prompt as @Image1, @Image2 and so on.

Which model for which photo: products, people, landscapes

The models differ less in whether they can animate a photo than in what they protect while they do it.

Our first pick per photo. Any model can win a given photo, so test before the final render.
The photoOpen firstWhy
A person's face, a family momentOpen firstKling V3 TurboWhyHolds a face well for the first eight to ten seconds, silent, and renders square prints at 1:1
A product on a surfaceOpen firstVeo 3.1WhyKeeps the photo's light and adds the sound of the thing: steam, a pour, a click
A pet mid actionOpen firstKling V3 TurboWhyFast natural motion, four legs that land; Seedance 2.0 Fast is the second try
A landscape or a cityscapeOpen firstWan 3.0WhyUp to 30 seconds for a slow drift, sound on by default; Gemini Omni Flash 1.1 for 4K
A car, a wedding, a danceOpen firstVeo 3.1WhyPhysics, reflections and two faces that hold, with the room's sound
A short loop for a feedOpen firstPika 2.2 or Luma Ray v3.2WhyShort silent clips, among the lowest cost per second here at 720p
A clip that must end on a second photoOpen firstSeedance 2.5WhyTakes a start and an end frame; so does Wan 3.0

The Kling page collects prompts for people and pets, the Veo page prompts that lean on sound and light, and the Seedance page the long takes and end frames.

Length, resolution and sound

Length is where the models part ways, and the model table further up lists every option. Veo 3.1 stops at 8 seconds, the Kling models at 15, and Seedance 2.5 and Wan 3.0 go to 30. A photo usually wants the short end: six seconds of a smile is a moment, fifteen is a loop.

Draft at 720p, or at 480p where it is offered, and render 1080p once the motion is right. On most models the bigger size costs more per second; Veo 3.1 prices 720p and 1080p the same. For 4K from a photo, Gemini Omni Flash 1.1 renders it natively, and so do Kling 3.0 4K and Kling O3 Omni 4K, both covered in our Veo 3.1 vs Kling 3.0 comparison.

Sound follows the model. Veo 3.1, Kling 3.0, Seedance 2.5, Seedance 2.0 Fast and Wan 3.0 render sound and have it on by default, with a switch to turn it off; on Veo 3.1 and Kling 3.0 a silent clip also costs less. Gemini Omni Flash 1.1 renders sound on every clip, and the model table marks the silent ones. For a photo of a person speaking, the better route is a talking avatar, which takes the portrait and an audio track and matches the mouth to the words.

Make the still first: generate, then animate

The best start frame is often one you made. A product shot that does not exist yet, a character you designed, a location you cannot visit: generate the still, judge it, then animate it. A still costs a fraction of a clip, so you pay for motion only once the frame is right.

  1. Open Nano Banana Pro and write the frame as a photograph: subject, setting, light, lens, and the clip's shape.
  2. Render two or three, and fix small things on the keeper with a photo edit rather than a re-roll: the image to image route takes the still and one line of change, and the Nano Banana page has frame prompts for it.
  3. Open the still as the start frame in the Video studio and describe only what moves.

Ten examples: six before-and-after pairs and four more to try

The first six are the real prompts behind the before-and-after pairs on our image to video page, which shows each photo beside the clip it became. The last four are new, written to the same rules and not rendered by us, so treat them as starting points. Use your own photo, keep the prompt, and change one thing at a time.

Six real pairs from our image to video page, and four new prompts written to the same rules.
The photoThe promptModel
A birthday cake, candles litThe promptShe leans forward and blows out the candles, the flames flicker and go out, smoke curls up, she laughs and claps, the hands around her clap too, handheld phone footageModelKling V3 Turbo, the page's clip
A dog on a beachThe promptThe dog drops the ball, crouches, then bolts along the beach through shallow water, sand and spray kicking up, the camera panning to followModelKling V3 Turbo, the page's clip
A coffee on a tableThe promptSlow push in as steam rises and curls from the coffee, sunlight shifting across the table, a hand enters and lifts the mugModelVeo 3.1, the page's clip
A couple at a weddingThe promptThey sway and turn slowly, her dress swinging, she laughs and leans her head on his shoulder, the string lights twinkling, guests clapping softly at the edges, handheldModelVeo 3.1, the page's clip
A parked carThe promptThe car pulls away along the coast road, palms passing, the sun flaring off the chrome, camera tracking from behindModelVeo 3.1, the page's clip
Kids by a sprinklerThe promptThey land and bounce again, hair flying, the sprinkler sweeping across the frame, one of them shrieking with laughter, handheldModelKling V3 Turbo, the page's clip
A grandfather reading by a windowThe promptHe looks up from the newspaper, takes off his glasses and smiles, the curtain moving in a breeze, the camera holds stillModelKling V3 Turbo, a new prompt
A perfume bottle on marbleThe promptA slow push in, light glints across the glass, a thin ribbon of mist drifts past, the bottle itself stays stillModelVeo 3.1, a new prompt
A mountain lake at dawnThe promptMist drifts across the water, the reflection ripples once, the clouds move slowly left, the camera holds stillModelWan 3.0, a new prompt
A city street at nightThe promptRain starts to fall, the neon signs flicker, the reflections on the wet road shimmer, a slow push inModelGemini Omni Flash 1.1, a new prompt

How to turn a photo into a video on DreamClip

  1. Open the Video studio with the start frame mode on: Kling V3 Turbo for people and pets, Veo 3.1 for products and sound, Seedance 2.5 when you need an end frame.
  2. Drop the photo on the start frame, cropped to the clip's shape.
  3. Write the motion line, one subject and one change, and pick the length and size.
  4. Run it. Credits are held when you press and settled only when a finished clip is delivered; an animation that fails is never charged.
  5. Judge the first second and the face. If either slips, calm the opening of the prompt or try the next model.

To choose between two models for one photo, run the same photo and motion line on each in the Video studio at its shortest length and smallest size, then render the keeper. Compare runs one text prompt on up to four video models side by side, each at its shortest length and smallest size in 16:9, with no charge for a failed lane; it takes no photo, so it picks a model for a motion, not for a picture.

Where image to video falls short

  • Hands and teeth. A hand that waves and a mouth that opens wide are where every model still slips. Keep the hands still, and use a talking avatar for speech.
  • A face turning past profile. The model has never seen the back of this head; it invents one. Keep turns under a quarter turn.
  • A second person entering. Anyone not in the photo is a stranger the model makes up. If they matter, put them in the still first.

Luma's Dream Machine, Pika's own app, Pollo and Veed all offer a photo to video step, as of October 2026, each inside its own editor or subscription. DreamClip runs Luma Ray v3.2 and Pika 2.2 beside Kling, Veo and Seedance in one studio, so the comparison is a model switch rather than a new account.

Questions

Is image to video AI free on DreamClip?

Trying it costs nothing up front. Signing up needs no card, and a free account gets credits every day. Image to video is priced by the model, the length and the size, so a three second clip at the smallest size on a lower cost model fits a day's credits far more easily than an eight second Veo 3.1 clip. A clip that fails is never charged.

Which image to video AI keeps faces the same?

Kling V3 Turbo and Kling 3.0 hold a face better than most for the first eight to ten seconds, and Veo 3.1 keeps the photo's light and colour over its 4 to 8 seconds. Every model drifts past ten seconds from one photo, so keep clips of people short.

Can I use my photo as the first frame on Veo 3.1?

Yes. Open Veo 3.1 in the Video studio with the start frame mode on, upload the photo and describe the motion. Veo 3.1 renders 4, 6 or 8 seconds at 720p or 1080p in 16:9 or 9:16, with sound on by default.

Can I set the last frame as well as the first?

On some models. Seedance 2.5, Seedance 2.0, Seedance 2.0 Fast, Wan 3.0, Wan 2.7, Gemini Omni Flash 1.1, MiniMax H3, MiniMax H3 Max and PixVerse V6 take a start and an end frame on DreamClip. Kling V3 Turbo, Kling 3.0 and Veo 3.1 take a start frame only.

Can I use Sora for image to video here?

No. DreamClip does not run Sora, and OpenAI retired the Sora API on September 24, 2026. For a photo to video clip, the models here are Kling V3 Turbo and Kling 3.0, Veo 3.1, Seedance 2.5 and the others in the table above.

Does Kling 4.0 do image to video on DreamClip?

No. Kling's own site announces Kling 4.0 (as of October 2026), and DreamClip does not run it. Every Kling model sold here, from Kling V3 Turbo to Kling O3 Omni 4K, takes a start frame and none takes an end frame; if the clip must end on a second photo, open Seedance 2.5 or Wan 3.0.

Related

Sources