The outfit post
A fit picture that already exists, turned into the three seconds of movement a feed rewards, without a second shoot or a tripod.
Full-body still in · A walk out · No rig, no keyframes
A stride, a coat that swings, a camera that backs away — built from a photo of someone standing perfectly still.
A real run on this page's model. The still on one side is the only input; the clip on the other is what came back. Watch the trench coat rather than the face — it swings a fraction behind the body, which is the detail a pan across a photo can never fake. Every generation is different, so treat this as the range rather than a result you will get back identically.

Drag to compare before and after
It takes one photograph of someone standing and generates a short clip of them walking, usually toward the lens. The person, the clothes and the location all come from your still; the stride, the fabric movement and the camera move are invented around them.
Three kinds of people search for this, and they want slightly different things. There is the person with one good outfit photo who needs three seconds of movement for a reel, because a still stops the scroll and a walk holds it. There is the illustrator or 3D artist with a character sheet, who wants to see the design move before committing to a rig.
And there is everyone with an old photograph of a relative standing in a doorway, who has wanted to see them walk out of it since the day they found it.
What separates a convincing result from an uncanny one is almost never the face — it is the feet. Walking is a balance problem: the body has to fall forward and catch itself, the planted foot has to stay planted while the frame moves past it, and the swinging arm has to oppose the leading leg.
Get any of that wrong and the brain flags it instantly even when it cannot say why. The other tell is the ground. If the pavement slides underneath someone who is not pushing off it, they are on a treadmill, and no amount of facial detail rescues the shot.
This page runs on Seedance 2.0 rather than the free engine, deliberately. Image-to-video across five seconds is exactly where the cheap tiers drift: the coat changes colour, the face loosens, the shoes become a different pair. Paying per second for a model that holds identity is the whole product here. Free credits cover a first run, the estimate updates with length and resolution before you commit, and a failed generation is refunded unless it broke the content policy.
Creative engine
Upload a still where the whole body is in frame, describe the walk and the camera, and check the estimate before you generate.
No skeleton, no motion capture, no keyframes — the walk cycle is generated from a single frame of someone standing.
Toward camera, away down a corridor, a slow strut or a hurried march: say it in the prompt and it changes the whole read.
Motion from a still runs through several tools here: two people instead of one in the AI hug video generator, a dance instead of a walk in the AI dance video generator, or carrying on from a clip you already have with the AI video extender.
Three decisions, and the first one decides most of the result.
Frame the whole body — hair to shoes — with a little floor visible below. A photo cropped at the waist has no lower body to build a stride from, and the model will invent one that does not match.
Toward the lens with the camera backing away, or away down a hallway with the camera following. Naming both movements separately is the difference between travel and a treadmill.
Short first: the gait either reads or it does not, and you will know inside two seconds. Once the stride looks right, rerun at the length you actually want.
One still, three very different reasons to make it move.
A fit picture that already exists, turned into the three seconds of movement a feed rewards, without a second shoot or a tripod.
See how a costume reads in motion — where the fabric catches, how the silhouette holds — before anyone spends a week rigging it.
A grandparent standing in a doorway in 1974, taking a few steps forward. It is the most-run version of this and the reason people find the tool at all.
What the generator does with a still, and where it struggles.
Poorly, and it is the single biggest cause of a bad result. If the frame stops at the waist the model has to invent a lower body from nothing, so the legs it builds rarely match the build, the trousers or the shoes you would expect. A full-length shot with some ground under the feet gives it everything it needs.
Because walking is a balance problem rather than a pose problem. The planted foot has to hold still against the ground while the frame moves past it; when the model loses that contact the feet skate and the whole stride reads as a treadmill. Shorter clips and a clearly described camera move both make it much less likely.
Yes, and stylised subjects often hold together better than photoreal humans do, because a viewer forgives an invented gait on a cartoon and forgives nothing on a real person. Character sheets and full-body renders work well as long as the whole figure is in frame.
Pick the length in the generator before you run. Longer is not automatically better here: a short clip that loops cleanly usually beats a long one that drifts halfway through, and since the model is billed by the second, the short version is also the cheap way to find out whether the gait works.
That is what the model on this page is chosen for. Identity across several seconds of motion is where cheaper video tiers give up first — the face loosens, the coat shifts colour, the shoes become a different pair. It is not guaranteed on every run, and a second attempt is usually the fix when it slips.
Only within reason. The location comes from your photo, so you can ask for a different camera move, a different pace or a different direction, but asking the model to relocate someone to a beach while also inventing a walk gives it two jobs and it will do both worse. Change the setting first, then animate the result.
The clip comes back silent by default, which is usually what you want — footsteps generated at the wrong tempo are more distracting than no footsteps at all. Add audio afterwards in whatever you edit in.
It is priced per second, and the estimate appears before you generate and updates with the length and resolution you choose. New accounts start with free credits, and a failed generation is refunded automatically unless it broke the content policy.
One credit pool covers Nano Banana images and Seedance and Veo video. The cost shows before every run, and failed jobs are not charged. Use a subscription for ongoing work, or a one-time pack when you just need to top up.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
One-time top-ups — buy extra credits any time you run low.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
Charges appear as “KAVEL AI” on your card statement. You can cancel any time from Settings → Billing; cancellation takes effect at the end of the period you already paid for.
Powered by
Use one balance across every supported image and video model — eleven of them today, listed below — and check the credit cost before each request.
Video models
Image models
Model credit guide
You're only charged for successful generations. The exact estimate in the generator varies by model, length, resolution, audio, and number of images.
| Type | Model | Credit cost |
|---|---|---|
| Video | Seedance 2.5 | The long-take tier, priced per second: 5s 480p ≈ 211 credits, 5s 720p ≈ 473. A full 30s take runs ≈ 1,260 at 480p and ≈ 2,835 at 720p. Supplying a reference clip lowers the per-second rate. |
| Video | Seedance 2.0 | 5s 720p image-to-video ≈ 188 credits; text-to-video ≈ 308 credits. Scales with resolution and length. |
| Video | Seedance 2 Fast | Faster and lower cost. 5s 720p text-to-video ≈ 248 credits. |
| Video | Seedance 2 Mini | The cheapest Seedance tier. 5s 720p ≈ 154 credits, 5s 480p ≈ 72 credits. |
| Video | Seedance 1.5 Pro | Audio doubles the rate. 5s 720p ≈ 27 credits silent, ≈ 53 with audio; 1080p ≈ 57 and ≈ 113. |
| Video | Veo 3.1 | Billed per video, not per second (Lite tier). About 45 credits at 720p and 53 at 1080p. |
| Video | Kling 3.0 | Audio raises the rate. 5s 720p ≈ 105 credits silent, ≈ 150 with audio; 1080p ≈ 135 and ≈ 203. |
| Video | MiniMax H3 | Fixed 2K, no resolution ladder. Priced per second — a 5s clip ≈ 158 credits. |
| Image | Nano Banana 2 | Generate or edit from text and images. About 8 credits per 1K image, 12 at 2K, 18 at 4K. |
| Image | Nano Banana Pro | Consistent run times across generations. About 12 credits per 1K or 2K image, 21 at 4K. |
| Image | Nano Banana 2 Lite | Faster, simpler variant. A flat 8 credits per image at every resolution. Start here to test. |
| Image | GPT Image 2 | The lowest-cost image model here. About 3 credits per 1K image, 6 at 2K, 12 at 4K. |
| Image | Seedream 5.0 Lite | Flat pricing — about 8 credits per image whether you generate or edit. |
One full-body shot, a described camera move, and a few seconds of someone walking who was standing perfectly still a minute ago.