FREE image generation with no sign-up, plus daily check-in credits that climb all week. No credit card, ever.
Sep 10, 2026

Most Realistic AI Video Model 2026: The Specs Don't Match

The most realistic AI video model you can run today is Veo 3.1. Google's own pages give three different answers for what it outputs, and OpenAI's Sora 2 announcement now ends with "The Sora product is no longer available" — which is why that one-line answer takes an article to establish.

Verdict: Veo 3.1. It is the only model in the table below that combines a third-party quality verdict, documented natively generated synchronized audio, and a published frame rate. The cost is an invisible SynthID watermark that cannot be switched off. PCMag UK rates Sora 2 as competitive with it until OpenAI's product change; the consumer Sora product is now retired.

Getting there is harder than it should be. Five of the models below have a documented conflict — an official page contradicting another official page, or contradicting the pages that quote it — and only three of thirteen publish a frame rate at all. Everything here was read on 10 September 2026. Six of these models — Veo 3.1, Kling 3.0, Seedance 2.5, LTX 2.5, Wan 3.0 and MiniMax H3 — run on Kavel from one balance.

Official output specifications for the AI video models compared here, read from vendor documentation on 10 September 2026

What each AI video model officially outputs

Most cells come from the vendor's own documentation, read on 10 September 2026; where the only published figure sits on a host or an API reference rather than the vendor's own page, the source column says which. Where nobody publishes a figure, the cell says so rather than borrowing one.

Model Vendor Official max resolution Frame rate Max duration Native audio Official source Read Conflict
Veo 3.1 Google 1080p and 4K on deepmind.google 24 fps 8s Yes deepmind.google 2026-09-10 Three Google pages, three answers
Sora 2 OpenAI 1280x720; 1920x1080 on sora-2-pro not stated 20s Yes developers.openai.com 2026-09-10 Consumer product retired
Kling 3.0 Kuaishou 4K per API update notice not stated 3–15s not stated kling.ai 2026-09-10 Third parties say 1080p, 10s
Seedance 2.5 ByteDance 1080p, 10-bit colour not stated 4–30s not stated seed.bytedance.com, docs.byteplus.com 2026-09-10 Third parties say 2K and 4K
LTX 2.5 Lightricks 4K not stated 20s not stated docs.ltx.io 2026-09-10 Internally consistent
Wan 3.0 Alibaba 480P, 1080P 30 fps 5s on the official help page Optional alibabacloud.com 2026-09-10 Hosts say 2–30s
MiniMax H3 MiniMax 480P, 768P not stated 5–15s on the vendor tool page Yes platform.minimax.io 2026-09-10 Doc compares H3 to itself
Vidu Q2 Vidu 540p, 720p, 1080p 24 fps 1–10s Yes platform.vidu.com 2026-09-10 Internally consistent
PixVerse V5 PixVerse 1080p not stated 5s; 1080p excludes 8s not stated aimlapi 2026-09-10 Resolution limits duration
Hedra Character-3 Hedra not stated for video — the vendor's 4K figure describes its image generator not stated not stated Audio is an input hedra.com 2026-09-10 Image spec reads as a video spec
HeyGen Avatar IV HeyGen "1280p+" on the product page; "4k, 1080p" in the API not stated Up to 30 min per the FAQ Yes, voice heygen.com 2026-09-10 "1280p" is not a resolution
Synthesia Studio Avatars Synthesia 3840x2160, stated as filming input not output not stated not stated not stated docs.synthesia.io 2026-09-10 Input mistaken for output
Digen 1 Turbo Digen 720p on the vendor site's model card not stated 5s not stated digen.ai 2026-09-10 Spec only on a model card, no doc page
Kavel Aggregator: runs Veo 3.1, Kling 3.0, Seedance 2.5, LTX 2.5, Wan 3.0, MiniMax H3 per model per model per model model pages 2026-09-10 Per-model options on each page

Only three of the thirteen AI video models publish a frame rate at all. LTX 2.5 and Vidu Q2 are the only two whose documentation does not contradict itself or its resellers.

Where the official AI video specs contradict each other

Where five AI video models' official specs disagree with a second official page or with the pages quoting them

Model One official source says Another source says The gap
Veo 3.1 deepmind.google: 1080p and 4K docs.cloud.google.com: 720p, 1080p 4K present or absent
MiniMax H3 platform.minimax.io: 480P, 768P third-party listings: 768p, 1080p, 2K three different ceilings
Wan 3.0 alibabacloud.com: duration 5s resellers: 2–30 seconds six-fold
Kling 3.0 kling.ai: 4K, 3–15s directories: 1920x1080, 10s both resolution and duration
Seedance 2.5 seed.bytedance.com: 1080p write-ups: 2K, then 4K upgraded in transit

Veo 3.1: three Google pages, three answers

Google DeepMind's model page promises "Professional grade resolution — Generate outputs in 1080p and 4K." The Gemini API documentation on ai.google.dev describes "8-second videos (720p, 1080p, or 4k)." The Gemini Enterprise Agent Platform page on docs.cloud.google.com lists "Supported output resolutions: 720p, 1080p" with no 4K at all, plus "Supported framerates: 24 FPS."

Three Google properties, three ceilings. In the one measurement we ran on 10 September 2026, the engine quoted deepmind.google and discarded the other two Google pages — worth knowing if you are comparing AI video specs from a summary rather than the source.

MiniMax H3: the doc compares the model to itself

MiniMax's video generation guide reads:

MiniMax H3: an open, general-purpose multimodal video model. It delivers mainstream 480P and 768P output and generates faster than MiniMax H3.

That sentence compares H3 to H3 — a copy-paste artefact sitting in the primary spec paragraph. The same page caps output at 768P while third-party listings advertise 768p, 1080p and 2K for the same AI video model.

Wan 3.0: five seconds or thirty

Alibaba Cloud's Model Studio help page states:

Resolution options: 480P, 1080P. Video duration: 5s. Defined specifications: 30 fps.

The hosts we checked describe "any duration between 2 and 30 seconds, with audio on or off."

A six-fold gap between a vendor's help page and the hosts reselling it is not rounding, and neither side says which configuration it is describing.

The Veo 3.1 model page on Kavel showing the duration and resolution options available at generation time

Kling 3.0: 4K in the changelog, 1080p in the directories

Kling's API update notice says:

Reference video duration has been expanded to 3–15 seconds, and 4K video generation is now supported.

Third-party directories list Kling 3.0 at 1920x1080 maximum; one guide states "a single Kling generation maxes out at 10 seconds; the default is 5 seconds."

Seedance 2.5: 1080p that grows in transit

ByteDance's Seed page describes a model that "can create 1080p videos with smooth motion, rich details, and cinematic aesthetics," and BytePlus documentation adds "up to 30 seconds" at "1080p (10-bit color depth)." That 10-bit detail is specific enough to be first-hand. In third-party write-ups the same AI video model becomes "1080p to 2K," then "a 30-second 4K clip at 24 FPS."

What actually drives realistic AI video

Realism is not one property. It decomposes into four things you can check before spending anything.

Native resolution, not upscaled output.
An AI video model that generates at 720p and exports a 1080p file is not producing 1080p detail. Vendors rarely separate the two, which is why resolution is the most contradicted row above.

Frame rate.
24 fps reads as film, 30 fps reads as video. Neither is inherently more realistic, but a mismatch with surrounding footage is immediately visible. Ten of the thirteen models here do not publish this number.

Whether the duration is one continuous shot.
A 30-second output assembled from cuts and a 30-second single pass are different products. Seedance 2.5 is the one model here described in a third-party write-up as generating its 30 seconds "in one continuous pass" — the vendor's own page does not say either way.

Whether audio is generated with the video or added afterwards.
Veo 3.1 and Sora 2 both describe natively generated, synchronized audio. Veo 3.1 and Sora 2 are the two models here whose documentation describes audio generated together with the video rather than added to it.

To compare two specific models on these axes instead of thirteen at once, each model page carries its current figures: Veo 3.1, Kling 3.0, Seedance 2.5, LTX 2.5, Wan 3.0, MiniMax H3.

Realistic people are a different category of AI video

General-purpose AI video models handle landscapes and product shots well and give themselves away on faces. A separate set of vendors builds only for the hardest case — a person talking — and their spec sheets have the same problem.

Hedra describes Character-3 as an "omnimodal character model combining image, text, audio, motion, and emotion for performance video." HeyGen's Avatar IV page is about turning photos into talking avatars. Synthesia's Studio Avatars documentation is written around filmed footage. Digen publishes its numbers on a model card rather than a documentation page: 5s at 720p. All four sit in the table above.

HeyGen's product page advertises:

Turn photos into talking avatars with 1280p+ video output

1280p is not a resolution. The standard ladder runs 720p, 1080p, 1440p, 2160p; HeyGen's own API reference lists the real options as "4k, 1080p". And Synthesia's 3840x2160 appears in filming guidelines as the resolution you should shoot at, which is exactly the kind of number that becomes an output claim once it has been copied twice.

Vendors building specifically for realistic people in AI video, and what their documentation actually states

Is your AI video watermarked, and can you sell it

Model Watermark Commercial use Source Read
Veo 3.1 Invisible SynthID, survives re-encoding, cannot be disabled Permitted, C2PA credentials referenced fal.ai, docs.cloud.google.com 2026-09-10
Sora 2 not stated in the API docs API terms apply; consumer product retired developers.openai.com 2026-09-10
Kling 3.0 not stated in the pages read Terms page located, not quoted here kling.ai 2026-09-10
Seedance 2.5 not stated not stated in the pages read docs.byteplus.com 2026-09-10
Wan 3.0 not stated not stated in the pages read alibabacloud.com 2026-09-10

Watermark and commercial-use terms for five AI video models, read from vendor documentation on 10 September 2026

The Veo row is the one to read twice: "All Veo 3.1 outputs are invisibly watermarked with Google's SynthID technology. The watermark persists through re-encoding and cannot be disabled." A visible watermark is a quality problem you can see. An embedded one is a provenance signal you cannot remove, and it ships with every realistic AI video that model produces.

Some of the watermark-free claims for other AI video models we came across sit on reseller pages rather than the model owner's. Read the terms on the platform you actually generate on, on the day you generate.

What each vendor admits its AI video still gets wrong

Vendor documentation is more candid than vendor marketing. PixVerse's API reference notes "the 1080p quality option does not support 8-second videos" — resolution and duration trade against each other, a constraint absent from every summary. Sora 2's own announcement page carries the retirement notice for the consumer product while the API continues.

None of the vendor pages we read publishes a blind comparison against a competitor. Third-party rankings exist — uk.pcmag.com names Veo 3.1 the quality winner among current AI video generators and describes Sora 2 as competitive until OpenAI's product change — and impressions circulate on scenith.in, invideo.io and huggingface.co. We do not run every model discussed here, so this article does not rank AI video by eye. A ranking assembled from clips we could not generate ourselves would not be worth reading.

The sharpest measurement of how realistic AI video has become comes from outside the review industry. The Internet Watch Foundation writes on iwf.org.uk that since it "first started monitoring AI in early 2023, we've seen a frightening advancement in the ability to generate" abuse material. That is a realism assessment from an organisation that reviews reported imagery, and it is the context in which non-optional provenance marks like SynthID make sense.

The AI video model shelf on Kavel — the six models from this comparison that run here, alongside the rest of the lineup

How to check an AI video model before you spend credits

Look for four numbers in the vendor's own documentation, not in a summary of it: native resolution, frame rate, maximum single-shot duration, and whether audio is generated with the video. Note the date you read them. Then generate one clip at the cheapest tier that answers your question before committing to a longer run. The pricing page shows what each tier costs in credits, and the fast variants exist for exactly this kind of test: Kling 3.0 Turbo, Seedance 2 Fast, Seedance 2 Mini.

FAQ

Which AI video model is the most realistic right now?
Veo 3.1, on the evidence available. It holds a third-party quality verdict from uk.pcmag.com, it is one of only two models documenting natively generated synchronized audio — the spec most visible when people speak on camera — and it publishes its frame rate, which most vendors do not. Sora 2 matched it on audio until OpenAI retired the consumer product. Run Veo 3.1 on its model page.

Why do sites list different resolutions for the same AI video model?
Because the vendors do. Google publishes three different maxima for Veo 3.1 across deepmind.google, ai.google.dev and docs.cloud.google.com, and Alibaba's help page differs from its resellers by a factor of six on Wan 3.0's duration.

Does higher resolution make AI video look more realistic?
Only if the model generates at that resolution instead of upscaling to it. Vendors rarely distinguish the two, which makes resolution the least reliable row in any AI video comparison.

Is AI video output watermarked?
Veo 3.1 output carries an invisible SynthID mark that survives re-encoding and cannot be turned off. Other vendors vary, and reseller claims about watermark-free output do not always match the model owner's terms.

What frame rate do these AI video models produce?
Three of the thirteen publish one: Veo 3.1 and Vidu Q2 at 24 fps, Wan 3.0 at 30 fps. The rest do not state it.

Where can I run these AI video models without six accounts?
All six of the models with model pages above run on Kavel from one balance — see the video model list and the pricing page.

Sources

Read 10 September 2026. Vendor documentation: deepmind.google, ai.google.dev, docs.cloud.google.com, aistudio.google.com, developers.openai.com, openai.com, kling.ai, ir.kuaishou.com, seed.bytedance.com, docs.byteplus.com, dreamina.capcut.com, docs.ltx.io, alibabacloud.com, platform.minimax.io, hailuoai.video, platform.vidu.com. Third-party pages consulted for the conflicts noted above: fal.ai, modelhunter.ai, atlascloud.ai, cined.com, segmind.com, docs.aimlapi.com, together.ai, scenith.in, creativebloq.com, invideo.io, huggingface.co, 3daistudio.com, admix.software, resource.digen.ai.

Last updated: September 2026.

Want to make your own AI video?

Turn an idea into a Kavel video in seconds. Pay only for what you use.