Kling 2.6 is the generation below Kling 3.0. It makes a 5 or 10 second clip from a prompt or a still image, and sound is a switch on the same run rather than a second step.
Two lengths, chosen up front. There is nothing in between and nothing longer.
Audio is generated with the shot, and it is what doubles the price.
Generated with Kling 2.6 on this site on 2026-08-04, silent, five seconds, unretouched. One clip rather than three: video is priced per second and a real example is worth more than a page of adjectives about one.
Kling 2.6 is the generation directly below Kling 3.0. It takes a prompt, or a still image plus a prompt describing what should happen to it, and returns a clip of either five or ten seconds. Two things about it are worth knowing before you use it. The first is that duration is an enum, not a number: five and ten are the only values the API accepts, so there is no trimming a seven-second idea down to fit a budget.
The second is sound. Audio is generated as part of the same run rather than added afterwards, which is genuinely better than dubbing something over a silent clip — and it costs about twice as much per second, so it is a decision rather than a default. There is no resolution choice at all on this generation; that arrived with Kling 3.0 Turbo, which trades the sound switch for a 720p/1080p one.
Creative engine
Per second, and the only thing that moves the rate is whether sound is on.
Older than Kling 3.0 and Kling 3.0 Turbo, and cheaper per second than either when run silent.
Five or ten seconds. The API rejects anything else.
Generated with the shot, not layered on, and priced accordingly.
Pick 2.6 when you want sound generated with the shot and you do not need 1080p. That is a narrower slot than it sounds, and it is a real one: Kling 3.0 Turbo, the obvious alternative, has no sound field at all, so a clip that needs audio has to come from 2.6 or 3.0. Run it silent while you are still finding the shot — at roughly half the rate, two silent attempts cost about what one sounded attempt does, and the motion is what you are testing anyway. Then turn sound on for the take you keep. Move up to Kling 3.0 when you need a resolution above what this generation produces or multi-shot sequences, and to 3.0 Turbo when speed matters more than audio.
The sound is made with the shot rather than dubbed over it afterwards.
Run silent, it is the lowest per-second rate of the Kling models offered here.
Hand it a photo and describe what should move, rather than describing a whole scene.
What this generation offers, and what the ones above it added.
Generate a clip from a written description with no source material.
Start from a still and describe the motion you want added to it.
A boolean on the request, priced at roughly double the silent rate.
Aspect ratio is chosen when generating from text; from an image it follows the source.
Taken from the published KIE (kie.ai) schema for this model, which is the API it runs on here.
Four steps, and the second one is where most clips are won or lost.
Starting from a still gives you control over what the scene looks like before anything moves.
What the camera does and what moves in frame are two different instructions. Merging them is how clips drift.
Half the rate and half the length, which is four times cheaper per attempt while you are still guessing.
Once the motion is right, run it again with audio rather than paying for sound on the tests.
Short clips where audio being native matters more than resolution.
Short vertical clips where the audio is part of the post rather than a music bed.
A still you already have, given a camera move and a little motion.
Rain, water, steam, crowds — scenes where the sound is most of the effect.
A silent five-second run is the cheapest way to find out if the motion works at all.
One credit pool covers Nano Banana images and Seedance and Veo video. The cost shows before every run, and failed jobs are not charged. Use a subscription for ongoing work, or a one-time pack when you just need to top up.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
One-time top-ups — buy extra credits any time you run low.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
Charges appear as “KAVEL AI” on your card statement. You can cancel any time from Settings → Billing; cancellation takes effect at the end of the period you already paid for.
Powered by
Use one balance across every supported image and video model — ten of them today, listed below — and check the credit cost before each request.
Video models
Image models
Model credit guide
You're only charged for successful generations. The exact estimate in the generator varies by model, length, resolution, audio, and number of images.
| Type | Model | Credit cost |
|---|---|---|
| Video | Seedance 2.0 | 5s 720p image-to-video ≈ 188 credits; text-to-video ≈ 308 credits. Scales with resolution and length. |
| Video | Seedance 2 Fast | Faster and lower cost. 5s 720p text-to-video ≈ 248 credits. |
| Video | Seedance 2 Mini | The cheapest Seedance tier. 5s 720p ≈ 154 credits, 5s 480p ≈ 72 credits. |
| Video | Seedance 1.5 Pro | Audio doubles the rate. 5s 720p ≈ 27 credits silent, ≈ 53 with audio; 1080p ≈ 57 and ≈ 113. |
| Video | Veo 3.1 | Billed per video, not per second (Lite tier). About 45 credits at 720p and 53 at 1080p. |
| Video | Kling 3.0 | Audio raises the rate. 5s 720p ≈ 105 credits silent, ≈ 150 with audio; 1080p ≈ 135 and ≈ 203. |
| Video | MiniMax H3 | Fixed 2K, no resolution ladder. Priced per second — a 5s clip ≈ 158 credits. |
| Image | Nano Banana 2 | Generate or edit from text and images. About 8 credits per 1K image, 12 at 2K, 18 at 4K. |
| Image | Nano Banana Pro | Consistent run times across generations. About 12 credits per 1K or 2K image, 21 at 4K. |
| Image | Nano Banana 2 Lite | Faster, simpler variant. A flat 8 credits per image at every resolution. Start here to test. |
| Image | GPT Image 2 | The lowest-cost image model here. About 3 credits per 1K image, 6 at 2K, 12 at 4K. |
| Image | Seedream 5.0 Lite | Flat pricing — about 8 credits per image whether you generate or edit. |
Common questions about this generation of Kling.
3.0 is the newer generation and adds a resolution choice and multi-shot sequences. 2.6 has no resolution field at all, and run silent it is cheaper per second than either 3.0 or 3.0 Turbo.
No. Duration is an enum of exactly 5 and 10 on this model, and anything else is rejected by the API.
About 83 credits for 5 seconds silent and about 165 with sound. Ten seconds is double each, so a sounded 10-second clip is roughly 330.
When the scene is about something audible — rain, a crowd, water — yes, because it is generated with the shot rather than dubbed over it. For a clip that will sit under a voiceover, no; you are paying double for audio you will mute.
Kling 3.0 Turbo, which takes a 720p or 1080p resolution. It has no sound option, which is the trade.
Credits are returned automatically, unless the request broke the content policy.
Describe the camera move and the subject move, keep it to five seconds and silent for the first attempt.
Start from a written description or hand it a still and describe the motion.