This is the tier built for scenes rather than shots: one call runs anywhere from 4 to 30 seconds, and the look, the motion and the pacing arrive through separate reference slots instead of all being crammed into the prompt.
One generation, any length from 4 to 30 seconds — no stitching between takes.
Images for the look, a clip for the motion, audio for the pacing, in the same request.
Every clip below is a single generation running the full thirty seconds — no cuts, no stitching, no second pass. What to watch is not the spectacle but what stays put across half a minute: the same face, the same wardrobe, the same room, the same light.
Seedance 2.5 is ByteDance's newest video model and the first tier here that generates a full scene in one pass. The Seedance 2 tiers are shot machines: describe a moment, get four to ten seconds of it, and anything longer becomes a stitching problem — matching light, matching wardrobe, matching the way a character walks, across clips that were never told about each other.
Most of the effort in AI video today goes into hiding that seam. Thirty seconds in a single generation removes it instead. The second change is what you are allowed to hand the model. One request accepts up to ten reference images, reference video clips and reference audio at the same time, and you can pin a first and a last frame and let it find the path between them.
Look, movement and rhythm each get their own slot rather than competing for room in one sentence of prompt. It runs at 480p or 720p in seven aspect ratios, billed per second, with generated audio as a switch rather than a separate job.
Creative engine
Rate is per second and depends on resolution and whether you supplied a source video.
The published schema takes any integer from 4 to 30, or -1 to let the model choose.
Not just stills — motion and pacing can be supplied as footage and a track.
Pin both ends of a shot when the start and finish are non-negotiable.
Pick 2.5 when the thing you need is longer than a shot, or when you already own the look. A thirty-second take that holds one character is worth more than four eight-second takes that nearly do, because the near-misses cost you an edit and usually a re-run. The same goes for references: if you have character sheets, a product photo, a colour key, or footage whose camera move you want copied, handing them over as references is more reliable than describing them and hoping. Where the Seedance 2 tiers still win is volume — they are cheaper per second, so prompt-finding and throwaway runs belong down there. The prompt transfers up.
One generation covers a beginning, a middle and an end, so there is no seam to hide.
References carry identity, motion and rhythm far more reliably than adjectives do.
Supplying reference footage lowers the per-second rate against generating motion outright.
What this tier accepts, and what comes out the other side.
Any duration in range in one generation, or -1 to let the model pick the length.
Character sheets, product shots or colour keys, all in the same request.
Copy a camera move from footage; drive pacing from a track. Reference footage totals up to 30s.
Sound produced with the picture, as a switch on the request rather than a second job.
1:1, 4:3, 3:4, 16:9, 9:16, 21:9 or adaptive — vertical comes straight out of the model.
Pin either end of a shot and let the model generate the path between them.
The generation parameters for this tier, from the provider's published schema.
Long generations reward different habits than short ones.
A five-second prompt describes a moment; a thirty-second prompt has to describe an order — what happens first, what the camera does when it changes, where it ends.
Look in the images, movement in a reference clip, pacing in reference audio. Prompts go vague when they carry all three.
Same prompt, same references, short duration. It is priced per second, so a test costs a sixth of the take.
Once the look is locked, re-run long. The full take is the deliverable; the short one was the rehearsal.
Jobs where the length and the references are the point.
Setup, turn and payoff inside one generation, rather than three clips and an edit.
Reference images hold identity for the full take, which is what short clips keep breaking.
Feed footage whose movement you want copied instead of describing the move in words.
Reference audio drives pacing, so moves and changes land on the beat.
One credit pool covers Nano Banana images and Seedance and Veo video. The cost shows before every run, and failed jobs are not charged. Use a subscription for ongoing work, or a one-time pack when you just need to top up.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
One-time top-ups — buy extra credits any time you run low.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
Charges appear as “KAVEL AI” on your card statement. You can cancel any time from Settings → Billing; cancellation takes effect at the end of the period you already paid for.
Powered by
Use one balance across every supported image and video model — eleven of them today, listed below — and check the credit cost before each request.
Video models
Image models
Model credit guide
You're only charged for successful generations. The exact estimate in the generator varies by model, length, resolution, audio, and number of images.
| Type | Model | Credit cost |
|---|---|---|
| Video | Seedance 2.5 | The long-take tier, priced per second: 5s 480p ≈ 211 credits, 5s 720p ≈ 473. A full 30s take runs ≈ 1,260 at 480p and ≈ 2,835 at 720p. Supplying a reference clip lowers the per-second rate. |
| Video | Seedance 2.0 | 5s 720p image-to-video ≈ 188 credits; text-to-video ≈ 308 credits. Scales with resolution and length. |
| Video | Seedance 2 Fast | Faster and lower cost. 5s 720p text-to-video ≈ 248 credits. |
| Video | Seedance 2 Mini | The cheapest Seedance tier. 5s 720p ≈ 154 credits, 5s 480p ≈ 72 credits. |
| Video | Seedance 1.5 Pro | Audio doubles the rate. 5s 720p ≈ 27 credits silent, ≈ 53 with audio; 1080p ≈ 57 and ≈ 113. |
| Video | Veo 3.1 | Billed per video, not per second (Lite tier). About 45 credits at 720p and 53 at 1080p. |
| Video | Kling 3.0 | Audio raises the rate. 5s 720p ≈ 105 credits silent, ≈ 150 with audio; 1080p ≈ 135 and ≈ 203. |
| Video | MiniMax H3 | Fixed 2K, no resolution ladder. Priced per second — a 5s clip ≈ 158 credits. |
| Image | Nano Banana 2 | Generate or edit from text and images. About 8 credits per 1K image, 12 at 2K, 18 at 4K. |
| Image | Nano Banana Pro | Consistent run times across generations. About 12 credits per 1K or 2K image, 21 at 4K. |
| Image | Nano Banana 2 Lite | Faster, simpler variant. A flat 8 credits per image at every resolution. Start here to test. |
| Image | GPT Image 2 | The lowest-cost image model here. About 3 credits per 1K image, 6 at 2K, 12 at 4K. |
| Image | Seedream 5.0 Lite | Flat pricing — about 8 credits per image whether you generate or edit. |
Answered from the published schema and our own runs.
Three things. Length: 2.5 generates 4 to 30 seconds in a single call, where the 2.0 tiers top out well short of that. Inputs: 2.5 takes reference video and reference audio alongside images, so motion and rhythm no longer have to be described in the prompt. Editing: it is built for changing part of an existing shot rather than only generating a new one. The 2.0 tiers stay on the site because they are cheaper per second and still suit short work.
Yes — one request, one output file, no stitching on our side. That is the whole point of the tier, and it is why the examples on this page are all exactly thirty seconds rather than a montage.
480p or 720p, in 1:1, 4:3, 3:4, 16:9, 9:16, 21:9 or adaptive, as mp4 or mov. Vertical 9:16 comes straight out of the model, so a social cut needs no reframing afterwards. There is no 1080p tier on this endpoint — for a higher-resolution finish, generate here and take the result through an upscale step.
Reference images set identity and look, and the model holds them across the full take — that is what makes a thirty-second shot of the same character possible. A reference clip supplies motion and camera language, so point it at footage whose movement you want, not whose content you want. Reference audio drives pacing, so a track with a clear beat produces moves that land on it. Total reference footage can run to thirty seconds.
Per second, with the estimate shown before you run it. A 5-second clip is about 211 credits at 480p and about 473 at 720p; a full 30-second take is about 1,260 and 2,835. Supplying a reference video lowers the per-second rate, which makes the sensible workflow the cheap one: test short, then commit to the length.
Long takes amplify a vague prompt. At five seconds an underspecified prompt simply picks something; at thirty it has time to drift — the camera wanders, or the scene resolves into a different one than you had in mind. The fix is sequence: say what happens in what order. Reference images tighten identity considerably, but they do not substitute for telling it where the shot is going.
Open the generator, hand it your references, and let it run the full thirty seconds.
Audio is generated with the picture rather than laid over it afterwards.