LTX's open-weights video model, running in your browser instead of on your own GPU. Text or image in, up to 4K and 20 seconds out, with a synchronised audio track — and the resolution you pick is the resolution you pay for.
Audio is generated with the picture, at no extra cost per second.
Long enough for a scene, not just a three-second loop.
The neon alley was generated on Kavel on 2026-08-17 at 720p — a real run through the same path this page's generator uses. The other two were shared on X by their creators, who state they made them with LTX 2.5: the 4K beetle by @todayskill_ai and the park scene by @what_the_func. Every clip here carries its own audio track.
The latest open-weights video model from LTX, the generative media company spun out of Lightricks. It generates synchronised audio and video together, natively up to 4K and 50fps, from either a text prompt or a still image.
Two things separate this model from the flagships it competes with. The first is that audio is not an add-on: sound is generated with the picture in the same pass, so a clip arrives finished rather than needing a track laid under it. The second is that the weights are open — published on Hugging Face with a day-one ComfyUI integration, and free to use for organisations under $10M in annual revenue.
That is why the clips people post are so often labelled with the GPU they ran on: this is a model you can host yourself if you have the hardware.
Most people do not have that hardware, which is what this page is for. Running LTX 2.5 here costs no setup: pick 720p, 1080p, 2K or 4K, pick a length between 2 and 20 seconds, and the credit estimate updates before you commit. A 6-second 720p clip took 23 seconds end to end on the run that produced the alley example above.
The resolution ladder is the steepest of any video model on Kavel — 4K costs eight times what 720p does per second — so the picker defaults to 720p rather than quietly starting you on the expensive rung.
Creative engine
Prompt or image in, 720p to 4K out, 2 to 20 seconds, audio included. Cost is shown before you generate.
Synchronised sound generated with the frames, included in the per-second price rather than charged as a premium.
Published on Hugging Face and integrated into ComfyUI on day one — run it on your own GPU, or run it here with none.
Pick it when the clip has to arrive finished, long, or sharp.
Ambience, footsteps, rain on metal — generated with the picture instead of sourced afterwards from a library that never quite matches.
Up to 20 seconds in one generation, which covers a whole beat rather than a fragment you have to loop.
Native 4K rather than an upscale, for anything going on a large screen or surviving a crop.
What the model accepts here, measured against the live API.
Start from a written prompt, or hand it a still and let it move — the first frame is yours to set.
720p, 1080p, 2K and 4K. The price moves with the rung, so the picker starts at 720p.
Length is a fixed ladder rather than a free number, and the longer rungs cost proportionally.
16:9 and 9:16, decided before generation so a vertical clip is composed vertically rather than cropped.
Read off the live API this model runs on here, on 2026-08-17.
Four habits, in the order they matter.
4K costs eight times as much per second. Find the shot you want cheaply, then re-run the keeper at the size you actually need.
A slow push, a static tripod, a handheld drift. Motion is the thing a still image cannot give you, so it is the thing worth spending words on.
Rain on metal, a room tone, distant traffic. The audio is generated from the same prompt, so leaving sound undescribed leaves it to chance.
Image-to-video pins the composition, palette and subject exactly, and lets the prompt spend all its words on movement instead.
Mostly work where a silent three-second loop would not have done the job.
Close work where 4K earns its cost — texture, water, small moving detail that falls apart at lower resolution.
9:16 composed as 9:16, with sound already attached, ready to post without a pass through an editor.
The 20-second ceiling covers a full atmospheric beat — a street, a landscape, a room waking up.
A photograph, a render or a generated frame, given a camera move and an ambient track.
One credit pool covers Nano Banana images and Seedance and Veo video. The cost shows before every run, and failed jobs are not charged. Use a subscription for ongoing work, or a one-time pack when you just need to top up.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
One-time top-ups — buy extra credits any time you run low.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
Charges appear as “KAVEL AI” on your card statement. You can cancel any time from Settings → Billing; cancellation takes effect at the end of the period you already paid for.
Powered by
Use one balance across every supported image and video model — eleven of them today, listed below — and check the credit cost before each request.
Video models
Image models
Model credit guide
You're only charged for successful generations. The exact estimate in the generator varies by model, length, resolution, audio, and number of images.
| Type | Model | Credit cost |
|---|---|---|
| Video | Seedance 2.5 | The long-take tier, priced per second: 5s 480p ≈ 211 credits, 5s 720p ≈ 473. A full 30s take runs ≈ 1,260 at 480p and ≈ 2,835 at 720p. Supplying a reference clip lowers the per-second rate. |
| Video | Seedance 2.0 | 5s 720p image-to-video ≈ 188 credits; text-to-video ≈ 308 credits. Scales with resolution and length. |
| Video | Seedance 2 Fast | Faster and lower cost. 5s 720p text-to-video ≈ 248 credits. |
| Video | Seedance 2 Mini | The cheapest Seedance tier. 5s 720p ≈ 154 credits, 5s 480p ≈ 72 credits. |
| Video | Seedance 1.5 Pro | Audio doubles the rate. 5s 720p ≈ 27 credits silent, ≈ 53 with audio; 1080p ≈ 57 and ≈ 113. |
| Video | Veo 3.1 | Billed per video, not per second (Lite tier). About 45 credits at 720p and 53 at 1080p. |
| Video | Kling 3.0 | Audio raises the rate. 5s 720p ≈ 105 credits silent, ≈ 150 with audio; 1080p ≈ 135 and ≈ 203. |
| Video | MiniMax H3 | Fixed 2K, no resolution ladder. Priced per second — a 5s clip ≈ 158 credits. |
| Image | Nano Banana 2 | Generate or edit from text and images. About 8 credits per 1K image, 12 at 2K, 18 at 4K. |
| Image | Nano Banana Pro | Consistent run times across generations. About 12 credits per 1K or 2K image, 21 at 4K. |
| Image | Nano Banana 2 Lite | Faster, simpler variant. A flat 8 credits per image at every resolution. Start here to test. |
| Image | GPT Image 2 | The lowest-cost image model here. About 3 credits per 1K image, 6 at 2K, 12 at 4K. |
| Image | Seedream 5.0 Lite | Flat pricing — about 8 credits per image whether you generate or edit. |
Common questions about generating with this model online.
You can start in the generator on this page, and every new account gets a credit balance to spend on it. Credits are shared across every model on Kavel, so nothing is locked to one of them. Video is the most expensive thing on the site to run, so a free balance goes further at 720p than at 4K.
Not here. The weights are open, so people do run it locally — the clips posted online are often labelled with the card they were rendered on. Running it on this page needs nothing but a browser; the hardware question becomes a credit question instead.
Yes, in the same pass as the picture rather than as a second step, and it is included in the per-second price. That is one of the model's defining features. Describe the sound you want in the prompt — leave it out and you still get audio, just not audio you chose.
54 credits for 6 seconds at 720p, 108 at 1080p, and 432 at 4K. The ladder is steep on purpose because the underlying cost is: 4K is eight times the per-second rate of 720p. The estimate is on screen before you generate.
Up to 20 seconds, chosen from a fixed ladder rather than typed as a free number. Longer clips cost proportionally more, so the usual approach is to find the shot at a short length and only extend the one that works.
Yes — image-to-video takes your still as the first frame and moves from there, which is the most reliable way to control exactly what the clip looks like. The prompt then only has to describe motion and sound rather than the entire scene.
Different shapes: this one sells a resolution ladder up to 4K and lengths to 20 seconds, while H3 runs a single flat rate with no resolution choice. Both are on Kavel, so the honest answer is to put one prompt through each — and there is a full side-by-side of the two in the blog if you want the numbers first.
Text or a still in, up to 4K and 20 seconds out, sound already attached. Start at 720p and only pay for size once the shot is right.
The community runs this on RTX 5090s. Here it runs in a browser tab.