Alibaba's image model for pictures that contain writing. Type a prompt with the exact words you want on the page, pick a ratio, and it comes back with those words spelled correctly — posters, infographics, storyboards, signage and UI mockups.
Headlines, body copy and captions, spelled the way you typed them.
Grids, numbered steps, panels and rules — not just a picture with a word on it.
Generated with the Pro tier on this site on 2026-08-05, unretouched, at 2K. Every character you can read in them came out of the model — a poster headline, a full page of infographic body copy, and six handwritten storyboard captions. Nothing was typeset afterwards.



This is the third generation of Alibaba's Qwen image line, and the thing it is built around is writing. Most image models treat text as decoration: ask for a shop sign and you get shapes that look like letters from a distance and fall apart up close. This model treats the words as content. Give it a headline, a set of numbered steps, a caption under every panel, and it lays the page out and spells the words.
That difference changes what you can ask for. A poster stops being a background you have to add type to in another tool, and becomes one generation. An infographic with four steps and a paragraph under each one becomes a prompt. The same capability covers newspapers, exam papers, menus, product packaging, game UI and interface mockups — anything where the information in the picture is the point of the picture.
It also renders more than one writing system, so the words do not have to be English. The tier that runs here is Pro, the top rung of the family, at 1K or 2K with a choice of seven aspect ratios.
Creative engine
Two resolutions, seven ratios. The ratio never changes the price.
Headlines, paragraphs and captions come back spelled, not suggested.
Grids, numbered sequences, panels and headers hold their structure.
Start from a prompt, or hand it a picture and describe the change.
Pick it the moment your prompt contains a word in quotation marks. If you are describing a scene — a fox in autumn woodland, a portrait in mixed light — a photoreal model like [Seedream 5.0 Pro](/image/seedream-5-pro) will serve you better, and [Nano Banana Pro](/image/nano-banana-pro) is the one to reach for when a design has to look art-directed. But the second the picture has to say something, the ranking inverts. A model that renders a beautiful café and puts MOFFEE CAFE over the door has not made you a usable image, and no amount of re-rolling fixes it reliably. This is the model that gets the sign right, and then gets the opening hours under it right too. It is also the model to use when the layout carries meaning: step one above step two, a header separated by a rule, a caption tied to the panel it belongs to. Those are structural decisions, and it makes them.
Posters, covers, flyers, packaging — where the type is the design.
Infographics, menus, exam papers, documents held in frame.
In-scene lettering that has to survive being looked at closely.
What the model accepts here, and what each control does.
A prompt in, a finished page out, with the wording carried through.
Hand it a picture and describe the change you want made to it.
1:1, 16:9, 9:16, 4:3, 3:4, 3:2 and 2:3 — portrait, landscape and square.
2K is the one to use when the smallest text on the page still has to read.
Measured against the live API this model runs on here, on 2026-08-05.
Four habits that decide whether the words come out right.
Write the exact string you want rendered — reading 'MARCH 14-22', not a description of it.
Headline at the top, a second line beneath, credits along the bottom edge.
Condensed grotesk capitals, handwritten script, a thin rule under the header.
Body copy and captions hold together at 2K in a way they cannot at 1K.
The pictures that have something to say.
A headline, a date line and a credit block, finished in one generation.
Numbered steps with real body copy under each one, laid out on a grid.
Panelled sheets where every frame carries its own caption.
Shopfronts, menus, packaging and interface screens with legible labels.
One credit pool covers Nano Banana images and Seedance and Veo video. The cost shows before every run, and failed jobs are not charged. Use a subscription for ongoing work, or a one-time pack when you just need to top up.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
One-time top-ups — buy extra credits any time you run low.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
Charges appear as “KAVEL AI” on your card statement. You can cancel any time from Settings → Billing; cancellation takes effect at the end of the period you already paid for.
Powered by
Use one balance across every supported image and video model — ten of them today, listed below — and check the credit cost before each request.
Video models
Image models
Model credit guide
You're only charged for successful generations. The exact estimate in the generator varies by model, length, resolution, audio, and number of images.
| Type | Model | Credit cost |
|---|---|---|
| Video | Seedance 2.0 | 5s 720p image-to-video ≈ 188 credits; text-to-video ≈ 308 credits. Scales with resolution and length. |
| Video | Seedance 2 Fast | Faster and lower cost. 5s 720p text-to-video ≈ 248 credits. |
| Video | Seedance 2 Mini | The cheapest Seedance tier. 5s 720p ≈ 154 credits, 5s 480p ≈ 72 credits. |
| Video | Seedance 1.5 Pro | Audio doubles the rate. 5s 720p ≈ 27 credits silent, ≈ 53 with audio; 1080p ≈ 57 and ≈ 113. |
| Video | Veo 3.1 | Billed per video, not per second (Lite tier). About 45 credits at 720p and 53 at 1080p. |
| Video | Kling 3.0 | Audio raises the rate. 5s 720p ≈ 105 credits silent, ≈ 150 with audio; 1080p ≈ 135 and ≈ 203. |
| Video | MiniMax H3 | Fixed 2K, no resolution ladder. Priced per second — a 5s clip ≈ 158 credits. |
| Image | Nano Banana 2 | Generate or edit from text and images. About 8 credits per 1K image, 12 at 2K, 18 at 4K. |
| Image | Nano Banana Pro | Consistent run times across generations. About 12 credits per 1K or 2K image, 21 at 4K. |
| Image | Nano Banana 2 Lite | Faster, simpler variant. A flat 8 credits per image at every resolution. Start here to test. |
| Image | GPT Image 2 | The lowest-cost image model here. About 3 credits per 1K image, 6 at 2K, 12 at 4K. |
| Image | Seedream 5.0 Lite | Flat pricing — about 8 credits per image whether you generate or edit. |
Common questions about generating with this model online.
You can start in the generator on this page without an account, and every new account gets a credit balance to spend on it. Credits are shared across every model on Kavel, so nothing is locked to one of them.
Yes. The tier wired here is the Pro one, the top rung of the family, which is the one that holds small text together. It runs at 1K or 2K.
The three examples on this page are the honest answer: a poster headline, four paragraphs of infographic body copy, and six handwritten captions, all generated here on 2026-08-05 and none of them typeset afterwards. Write the exact wording into your prompt and it renders that wording.
Yes — multilingual rendering is one of the things the 3.0 generation was built for. Put the exact characters you want in the prompt.
About 16 credits at 1K and about 30 at 2K. The aspect ratio does not change either number.
Yes. Upload a source image in the generator and describe the change — the same text-rendering strength applies when you are adding or replacing writing in a picture you already have.
Write the words you want on the page, pick a ratio, and let it typeset the thing for you.
Run your first prompt in the box above without an account.