The remote presenter
A talk delivered from a spare room reads very differently against a plate that matches your seated height and your desk lamp. It is the cheapest production upgrade available to anyone working from home.
One description · Locked-off framing · Centre left empty
Stock backdrops are the reason so many keyed shots look wrong: the camera is at the wrong height, the light comes from the wrong side, and forty other channels are using the same file. Describe the room you want to be standing in instead.
Both generated on this page's engine from a single description each. They are shown side by side because a plate has no before — what matters is whether each one leaves you somewhere to stand. Every output is AI-generated, so yours will differ.

Asked for a curved news-desk wall with cove lighting and a blurred skyline panel, framed from a locked-off camera at seated height. Notice what is missing: no desk in the middle, no chair, no presenter. The interesting detail — the lit cove, the reflective floor — all sits at the edges, and the middle third is deliberately dull so a keyed subject has somewhere to be.
Take the studio
Golden hour, a low parapet across the lower third, unbranded towers behind. The parapet is doing real work: it gives a horizon line for your feet to relate to, which is what stops a standing subject looking like a sticker. The sun sits off to one side rather than behind the centre, so you can light yourself to match without fighting a halo.
Take the rooftopKeying is easy now. Matching is what still gives people away.
Modern software will pull a clean key off a decently lit green screen in seconds, and that has quietly moved the difficulty somewhere else. What now separates a convincing composite from an obvious one is agreement between two images that were never in the same room: eye level, lens character, light direction, colour temperature, and depth of field.
Get those wrong and the key can be perfect while the shot still reads as a person pasted onto a poster.
That is why a generated plate has an advantage over a stock one that has nothing to do with novelty. You can specify the conditions. Seated eye level for a podcast, standing for a presentation. Soft even light from the left because that is where your key light is. Shallow focus so the background falls away the way a long lens would render it.
A locked-off camera, because a plate with implied motion in it fights a camera that is not moving. None of that is available when you are picking from a library.
The honest boundaries: what you get is a still image, so it holds the shot but nothing in it moves — no drifting clouds, no traffic, no flicker on the screens. It also does not key your footage for you; that is your streaming software or editor's job, and the screen behind you still has to be lit evenly for it to work at all.
If what you actually want is the background of an existing photo cleaned up rather than replaced, the product photo editor does that job in place.
Creative engine
Describe the room, the eye level and the light, then key your footage onto it in your own software.
Seated, standing, or low. It is the one parameter that ruins a composite on its own, and the one people never think to specify.
Look at which side your key light is on and ask for a plate lit from the same side. Two contradicting light directions read as fake instantly.
A plate description is a camera report as much as a scene description.
"Shot from a locked-off camera at seated eye level" or "at standing eye level". This is the parameter that decides whether you look like you are in the room or in front of it.
Soft window light from the left, hard overhead, golden hour from camera right. Then set your own key light to agree with it. Two lights disagreeing is the second commonest tell after eye level.
Say it explicitly. Every instinct an image model has is to fill the middle of the frame with the subject, and you need that space. Push the detail to the edges.
No lettering, no logos, no other figures. Trademarks in a background are a real problem on a monetised channel, and a stray person in the plate will appear to walk through you.
Four situations where the library file is the thing holding the shot back.
A talk delivered from a spare room reads very differently against a plate that matches your seated height and your desk lamp. It is the cheapest production upgrade available to anyone working from home.
The background is the part of a stream you cannot renovate. Swapping in a rooftop or a studio is a set change that costs nothing and does not require moving furniture, next to the avatar as the other way to control what viewers see.
A consistent plate across forty lessons recorded over six months hides the fact that your room changed, the season changed and you moved house halfway through.
Two people in two cities look like one conversation when both plates are described from the same brief — same eye level, same light direction, same depth of field.
What people check before keying onto a generated image.
No — it makes the image you key onto. The keying itself happens in OBS, Zoom, Premiere, Resolve or whatever you already use, and it still depends on lighting your actual green screen evenly. This page removes the background problem, not the lighting one.
Not always. Most streaming software now offers a backgroundless mode that segments you without one, and these plates work fine behind that. A real physical screen still gives a cleaner edge around hair and glasses, which is why anything recorded rather than streamed usually still uses one.
The plate comes back at a size suited to a 16:9 frame, which covers a 1080p stream comfortably. For a 4K timeline, generate at the largest option available and check the edges at full size before you build a series around it — an upscaled plate shows its softness fastest in fine architectural detail.
Not from this page — what you get is a still, and a still is genuinely the right choice for most talking-head work because moving backgrounds pull attention off the speaker. If you do want motion, take the plate you like and run it through image-to-video for a slow ambient loop.
Yes, including monetised streams, client work and course material. The reason the prompts here exclude signage and logos is precisely that: a plate with a real brand visible in it is a liability regardless of who generated it.
Work through it in this order: eye level, then light direction, then colour temperature, then focus. Nine times in ten it is the first two. A plate shot from standing height behind a seated presenter cannot be fixed by grading — regenerate it at the right height instead.
One credit pool covers Nano Banana images and Seedance and Veo video. The cost shows before every run, and failed jobs are not charged. Use a subscription for ongoing work, or a one-time pack when you just need to top up.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
One-time top-ups — buy extra credits any time you run low.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
Charges appear as “KAVEL AI” on your card statement. You can cancel any time from Settings → Billing; cancellation takes effect at the end of the period you already paid for. Operator details, the full model list, and the refund window are on the about page.
Powered by
Use one balance across every supported image and video model — all of them listed below — and check the credit cost before each request.
Video models
Image models
Model credit guide
You're only charged for successful generations. The exact estimate in the generator varies by model, length, resolution, audio, and number of images.
| Type | Model | Credit cost |
|---|---|---|
| Video | Seedance 2.5 | The long-take tier, priced per second: 5s 480p ≈ 211 credits, 5s 720p ≈ 473. A full 30s take runs ≈ 1,260 at 480p and ≈ 2,835 at 720p. Supplying a reference clip lowers the per-second rate. |
| Video | Seedance 2.0 | 5s 720p image-to-video ≈ 188 credits; text-to-video ≈ 308 credits. Scales with resolution and length. |
| Video | Seedance 2 Fast | Faster and lower cost. 5s 720p text-to-video ≈ 248 credits. |
| Video | Seedance 2 Mini | The cheapest Seedance tier. 5s 720p ≈ 154 credits, 5s 480p ≈ 72 credits. |
| Video | Seedance 1.5 Pro | Audio doubles the rate. 5s 720p ≈ 27 credits silent, ≈ 53 with audio; 1080p ≈ 57 and ≈ 113. |
| Video | Veo 3.1 | Billed per video, not per second (Lite tier). About 45 credits at 720p and 53 at 1080p. |
| Video | Kling 3.0 | Audio raises the rate. 5s 720p ≈ 105 credits silent, ≈ 150 with audio; 1080p ≈ 135 and ≈ 203. |
| Video | MiniMax H3 | Fixed 2K, no resolution ladder. Priced per second — a 5s clip ≈ 158 credits. |
| Image | Nano Banana 2 | Generate or edit from text and images. 20 credits per 1K image, 30 at 2K, 45 at 4K. |
| Image | Nano Banana Pro | Consistent run times across generations. 30 credits per 1K or 2K image, 50 at 4K. |
| Image | Nano Banana 2 Lite | Faster, simpler variant, and the everyday editing price: a flat 8 credits per image at every resolution. |
| Image | GPT Image 2 | The lowest-cost premium image model here. 10 credits per 1K image, 15 at 2K, 30 at 4K. |
| Image | Seedream 5.0 Lite | Flat pricing — 20 credits per image whether you generate or edit. |
Describe the room, the eye level and the light, then key onto it in the software you already use — or put an avatar there instead.