Seedance 2 makes video and sound in a single generation. Turn text or images into moving video, then steer the result with start and end frames or with image, video, and audio references. Output runs 480p to 1080p at 4 to 10 seconds, and the credit cost is shown before you generate.
Generate sound that matches the shot in the same run (optional).
Guide the result with start/end frames and multimodal references.
Video with audio, from text or from an image. Real examples generated with the model.
Seedance 2 is ByteDance's video model that turns text or images into moving video. Unlike tools that render silent clips, it can produce shots with dialogue and lip sync and generate the audio in the same pass, so a talking scene comes out finished instead of needing a second sound step. Alongside text-to-video, the model supports start and end frames plus multimodal references built from images, video, and audio, so you can steer a shot toward what you pictured rather than re-rolling the prompt until it lands.
Here it runs 480p to 1080p at 4 to 10 seconds, which covers everything from a quick social clip to a polished hero cut; the model itself is documented up to 4K. On Kavel, Seedance runs, the credit estimate updates before you commit, and failed generations are refunded automatically (unless it broke the content policy).
It sits alongside faster, lower-cost variants, so you can test an idea cheaply and switch to the full model only when a clip earns the higher quality and resolution.
Creative engine
Choose an input method, set resolution, length, and audio, and the credit estimate updates before you generate.
Describe the subject, motion, camera, and mood in one prompt, and the model turns that description into a moving shot with matching pacing.
Start from a text prompt, from start and end frames, or from image, video, and audio references when you want tighter control over the result.
Turn audio on and the model generates sound that matches the shot in the same run, so a spoken or musical clip finishes in one pass.
The tiers here charge more per second the higher they go. When a shot only needs one quality level and picking a tier is just overhead, MiniMax H3 returns a fixed 2K and bills a single flat rate.
Continuous full-body motion is where this model separates from ones built mainly for short camera moves — a dance is a single unbroken action, and a model that quietly resets between beats gives you limbs that jump. If that is the shot you want, the AI dance video generator starts from one photo with the framing already set up for it.
Reach for this model when you want video and sound in one generation instead of two, and when reference material should steer the shot rather than luck. The three variants let you trade quality for speed and cost without changing your workflow. That flexibility matters most when a project mixes quick social cuts with a few hero shots, because the same prompt and the same references carry across every variant without a change to your workflow.
The model generates matching audio in the same pass, so dialogue and lip-sync clips need no separate step.
Beyond start and end frames, it accepts image, video, and audio references to direct the result.
Use the full model for quality, or the Fast and Mini variants when speed or cost matters most.
From text-to-video to native audio, the model covers the parts of a shot that usually take separate tools.
Turn subject, motion, camera, and mood into video from one prompt.
Pin a still to the first or last frame to lock the motion.
Pass image, video, and audio as references to direct the shot.
Enable generate_audio and the model adds sound that fits the shot.
Render at 480p, 720p or 1080p to match the use.
The generation parameters for this model, from the provider's published spec.
Four steps from prompt to a finished clip.
Start from text-to-video, or from start/end frames and references.
Describe the subject, action, camera, and style.
Choose the resolution, clip length, and whether to generate audio.
Confirm the estimate, run it, and download the result.
From short-form clips to talking scenes, here is where the model earns its place in a workflow.
Make scroll-stopping short clips for posts, ads, and stories, with sound baked in.
Generate audio in the same pass to make clips where people actually speak.
Turn key visuals and product shots into moving video that shows the item in motion.
Chain reference frames to bring a storyboard to life without a crew, a camera, or a location shoot.
One credit pool covers Nano Banana images and Seedance and Veo video. The cost shows before every run, and failed jobs are not charged. Use a subscription for ongoing work, or a one-time pack when you just need to top up.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
One-time top-ups — buy extra credits any time you run low.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
Charges appear as “KAVEL AI” on your card statement. You can cancel any time from Settings → Billing; cancellation takes effect at the end of the period you already paid for.
Powered by
Use one balance across every supported image and video model — eleven of them today, listed below — and check the credit cost before each request.
Video models
Image models
Model credit guide
You're only charged for successful generations. The exact estimate in the generator varies by model, length, resolution, audio, and number of images.
| Type | Model | Credit cost |
|---|---|---|
| Video | Seedance 2.5 | The long-take tier, priced per second: 5s 480p ≈ 211 credits, 5s 720p ≈ 473. A full 30s take runs ≈ 1,260 at 480p and ≈ 2,835 at 720p. Supplying a reference clip lowers the per-second rate. |
| Video | Seedance 2.0 | 5s 720p image-to-video ≈ 188 credits; text-to-video ≈ 308 credits. Scales with resolution and length. |
| Video | Seedance 2 Fast | Faster and lower cost. 5s 720p text-to-video ≈ 248 credits. |
| Video | Seedance 2 Mini | The cheapest Seedance tier. 5s 720p ≈ 154 credits, 5s 480p ≈ 72 credits. |
| Video | Seedance 1.5 Pro | Audio doubles the rate. 5s 720p ≈ 27 credits silent, ≈ 53 with audio; 1080p ≈ 57 and ≈ 113. |
| Video | Veo 3.1 | Billed per video, not per second (Lite tier). About 45 credits at 720p and 53 at 1080p. |
| Video | Kling 3.0 | Audio raises the rate. 5s 720p ≈ 105 credits silent, ≈ 150 with audio; 1080p ≈ 135 and ≈ 203. |
| Video | MiniMax H3 | Fixed 2K, no resolution ladder. Priced per second — a 5s clip ≈ 158 credits. |
| Image | Nano Banana 2 | Generate or edit from text and images. 20 credits per 1K image, 30 at 2K, 45 at 4K. |
| Image | Nano Banana Pro | Consistent run times across generations. 30 credits per 1K or 2K image, 50 at 4K. |
| Image | Nano Banana 2 Lite | Faster, simpler variant, and the everyday editing price: a flat 8 credits per image at every resolution. |
| Image | GPT Image 2 | The lowest-cost premium image model here. 10 credits per 1K image, 15 at 2K, 30 at 4K. |
| Image | Seedream 5.0 Lite | Flat pricing — 20 credits per image whether you generate or edit. |
Common questions about the Seedance 2 AI video generator.
Seedance 2 is ByteDance's video model that turns text or images into video. It supports start and end frames plus image, video, and audio references, and it can generate audio in the same pass. Output here runs 480p–1080p at 4 to 10 seconds.
Yes. Turn on generate_audio and the model adds sound that matches the shot in the same generation, which suits dialogue and lip-sync clips without a second step.
Credits depend on resolution, length, and input. For example, 720p at 5 seconds is about 308 credits. The estimate appears before you generate, and failed generations are free.
Yes. Pass start and end frames or image and video references, and the model generates moving video from them, which is the fastest way to animate a still you already have.
Describe the subject, motion, camera work, and mood in detail. Adding references keeps the composition and flow stable across the clip.
Fast and Mini are faster, lower-cost versions of the model. Use the full Seedance 2 for quality, or Fast and Mini when speed or cost matters most.
Fast is for iterating: same family, lower cost, quicker turnaround. Switch to 2.0 for the final render — it holds detail and motion better at 720p and 1080p, which is exactly the shot you keep.
Open the generator, pass text or an image, check the credits, and generate.
Pick the resolution and length that fit your use.