Three different kinds of software turn a face photo into a talking video, and one question tells you which you want: does the talking photo need to lip sync to your own audio recording, or is a generated voice fine? Answer that and eleven tools collapse into three routes. We read all eleven talking photo pages on 2026-09-11 and ran one of them on a portrait we made ourselves, which is how we found the thing none of them mention.
Have an audio file? Use an avatar tool — HeyGen, Hedra, TalkPix and lip-sync.ai all accept an uploaded recording. Just want a photo to say one line, voice unimportant? A video model does the whole talking photo in one pass: picture, voice and room tone generated together. On Kavel that is Veo 3.1 at 115 credits for a 720p clip, and the 40 credits a new account starts with cover it. Every number here was read on 2026-09-11.

Three routes that turn a face photo into a talking video
Search results mix all three together. They are three different products, and a talking photo made on each one fails in a different way.
Route 1 — Avatar tools: photo, script, and a voice
You upload a portrait, type what the talking photo should say, pick a synthetic voice, and the service renders it. HeyGen, Hedra, Vidnoz, TalkPix, DomoAI and lip-sync.ai work this way. Most also accept an audio file instead of a script, which is the detail that matters: HeyGen's lip sync page says you can upload an audio file or record your own voice, and Hedra's page lists three ways to supply it — upload a recording, record directly, or type a script and pick a voice (both read that day).
Route 2 — Video models: photo and one line of direction
You upload the portrait, describe what happens, and the model renders the whole talking photo with its own generated speech and ambient sound in one pass. Veo 3.1, Seedance 2.0 and Kling 3.0 work this way. No voice picker, no audio upload — the model writes the performance. Of the six pages we read that the answer engines currently cite, four do not mention this route at all.
Route 3 — Local and open source: nothing leaves your machine
SadTalker describes itself as "Audio-Driven Single Image Talking Face Animation" and Wav2Lip as speech-to-lip generation — both take a still and an audio file (repos read 2026-09-11). LivePortrait belongs here too but works differently: its repo calls it portrait animation, and it is driven by a video rather than audio, so it moves a face rather than making it speak. Nothing is uploaded to anyone and there is no subscription. The cost is a working Python environment and your own hardware.

Talking photo pricing, free plans, watermarks, clip length and export resolution
| Tool | Route | Free tier | Watermark on free | Length per clip | Upload your own audio | Paid entry | Read |
|---|---|---|---|---|---|---|---|
| HeyGen | Avatar SaaS | 3 videos / month | Yes, removal is a Creator feature | 1 min free, 30 min paid | Yes, audio file or recording | $29 / mo | 2026-09-11 |
| Hedra | Avatar SaaS | Page title says free; plan list starts at $15 | Not stated on pricing page | Not stated on pricing page | Yes, upload, record or type | $15 / mo, 1,500 credits | 2026-09-11 |
| Vidnoz | Avatar SaaS | 5 photo avatars / day | Yes on the free tier | Not stated | Voiceover upload is a plan row | $19.99 / mo (second-hand) | 2026-09-11 |
| TalkPix | Avatar SaaS | None, packs only | No, paid exports are clean | 7s example, 1 credit / second | Yes, script or audio | $5 pack, no subscription | 2026-09-11 |
| DomoAI | Avatar SaaS | None listed | No, watermark-free outputs | 5 / 10 / 20s, 60s on Pro | Tutorials say yes; not on its pricing page | $29 / mo billed yearly | 2026-09-11 |
| VEED | Editor | Yes | Yes (second-hand) | 10 min export cap (second-hand) | Yes, it is an editor | $9 / mo | 2026-09-11 |
| lip-sync.ai | Avatar SaaS | Yes, no sign-up | Not stated | Not stated | Yes, audio-driven | Not stated | 2026-09-11 |
| Kavel · Veo 3.1 Lite | Video model | 40 credits at sign-up + daily check-in | No | 8s at 1080p | No, the model writes the voice | 115 credits at 720p | 2026-09-11 |
| Kavel · Seedance 2.0 | Video model | Same shared balance | No | 4–10s, up to 1080p | No, audio in the same pass | 155 credits / second at 720p | 2026-09-11 |
| Kling 3.0 | Video model | Not stated on the plan page | Watermark wording on the plan page | Not stated | No, sound is generated | Not stated | 2026-09-11 |
| SadTalker · Wav2Lip | Local / open source | Free, runs on your machine | No | Length of your audio | Yes, audio is the input | $0 + your own hardware | 2026-09-11 |
| LivePortrait | Local / open source | Free, runs on your machine | No | Length of the driving video | No, driven by video, not audio | $0 + your own hardware | 2026-09-11 |

Second-hand means the figure came from a third party because the vendor's own page would not render without JavaScript. It is marked, not hidden.
What each free talking photo plan actually gives you
-
HeyGen, read that day: Free is $0/mo for 3 videos a month, 1 minute each, with Avatar IV access and 30+ languages. Watermark removal is a Creator feature, so free output carries one. Creator is $29/mo: 600 credits, 30-minute videos, 1080p, unlimited photo avatars, voice cloning. Pro is $49/mo for 1,000 credits and 4K.
-
Hedra, read the same day: the page is titled "Free Plan, No Credit Card" while the plans listed start at Basic $15/mo for 1,500 credits with slower generations. Creator is $30/mo for 5,400 credits, Professional $75/mo for 14,400. Every tier says commercial use. No watermark policy and no clip length are stated.
-
Vidnoz, read the same day: free gets 5 photo avatars a day against 100/day and unlimited above it, and "Voiceover Uploads" is its own row in the matrix. Its own talking photo page says the free version may include watermarks.
The two that do not sell a subscription
-
TalkPix, read the same day: no free tier, no subscription, packs from $5, billed 1 credit per second — its own 7-second 720p sample costs 7 credits. Paid exports are clean, and its input line is explicit: "One portrait + a script or audio."
-
DomoAI, read the same day: Standard is $29/mo billed annually for 2,200 credits with Talking Avatar at 5s, 10s and 20s; Pro is $99/mo for 8,000 credits and up to 60s. Outputs are watermark-free.
-
lip-sync.ai, read the same day: free, audio-driven, no sign-up, and no published limit, watermark policy or price.
Talking photo text to speech, uploading your own audio, and voice cloning
This is the split. If the clip needs a specific voice (yours, a client's, a recorded interview), routes 1 and 3 are the only answers, and the audio column above is the one to read.
Without a recording the calculation inverts. On the avatar route you now have to choose a synthetic voice, and a wrong choice is what makes a talking photo feel like a corporate explainer. A video model skips the step: it writes the delivery along with the picture. HeyGen offers voice cloning on Creator and 175+ languages; Vidnoz lists voice cloning as a paid add-on. Neither Veo 3.1 nor Seedance 2.0 offers a voice picker at all.

What makes a talking photo fail: the portrait, the light and the resolution
The input these tools ask for is consistent across their own help pages: one face, facing the camera, evenly lit, in focus, mouth closed and visible. Side profiles, sunglasses, heavy shadow across the mouth and two faces in frame are the conditions their guidance tells you to avoid.
There is a fifth failure mode none of that guidance mentions, because it is not about your photo being bad. On the video-model route the face can come back subtly altered even when the input is perfect. We only caught it because we kept the original file.
We ran it ourselves: one photo in, one talking photo out
We generated a portrait so no real person's face was involved, uploaded that still, and ran Veo 3.1 Lite at 1080p on Kavel on 2026-09-11. The prompt asked for one spoken sentence and, as explicitly as we could put it, for nothing else to change: "Keep her face, hairline, skin texture, jumper and the plain grey wall behind her exactly as they are in the source photograph — do not restyle, beautify, relight or change her identity."
What came back was a talking photo of 8 seconds, 1920x1080, H.264 with a real AAC track (mean -19.6 dB, peaks at -2.5 dB — speech, not silence). The mouth tracks the sentence properly, which is the one thing a talking photo has to get right.

As a talking photo it works. Two things changed that we did not ask for.
The output is 16:9 and our still was 4:5, so the model re-framed and pushed in, losing the shoulders and the jumper. That one is ours — we left the aspect ratio at its default.
The second change is not ours.

Side by side, the eyebrows are heavier, the skin is smoother, the jaw is narrower, and a neutral closed mouth has become a slight smile. It is close. On a phone, at a glance, you would not stop. It is not the same face — and the prompt said not to beautify.
We can say that only because we still have the file we uploaded. That is the whole trick and it costs nothing: keep the original. Nearly every talking photo comparison shows you one talking photo output and asks you to agree it looks good. An output on its own cannot tell you whether it is still the same person.
If identity has to survive, whether that is a client, an employee or anyone recognisable, check the output against the file you sent, whichever route you use. We ran this test on one video model and one photo; we have not run the same test on the avatar tools, so treat this as a reason to check rather than a ranking.
What the talking photo pages the answer engines already cite leave out
We read the pages ChatGPT and Perplexity currently cite for this question on 2026-09-11: veed.io's talking-head API guide, techsifted.com's avatar roundup, aipedia.wiki's avatar guide, en.ai-pedias.com's talking photo piece, pikvue.com and comparegen.ai's Synthesia-vs-HeyGen-vs-D-ID comparisons, flowjam.com's how-to, teachustechnology.com's D-ID walkthrough, capcutguide.com's Pippit pricing page, and dreamina.capcut.com's own image-to-video page.
Two things stand out. Not one of them covers all three routes. veed.io covers avatar tools and the open-source models but not the video models; aipedia.wiki covers avatar tools and video models but not open source; the remaining four we counted cover avatar tools only. And across the six we measured there are 114 dollar figures between them, against at most three dated-reading markers on any single page. Two of the six have none at all. A talking photo price with no date on it is a guess about a page that has probably already changed.

That is why every cell in the table above carries a date.
The three brands the answer engines check that are not talking photo tools
Ask an assistant for a talking photo tool and it goes and checks five vendor sites. Three of them do not do what you asked:
- Synthesia leads with AI avatars and 1,000+ AI voices in 160+ languages, aimed at corporate video (its own site, read 2026-09-11). It is a strong product, but a single uploaded snapshot is not what it is built around.
- D-ID's own navigation now opens with Agents, AI Avatars and Agentic Videos and calls itself a digital human platform (read 2026-09-11). The photo-to-video entry is still there; it is no longer the front of the product.
- Captions describes itself as AI that "edits like a professional editor would" (read 2026-09-11): a video editor, not a photo-driven generator.
If the subject is a pet rather than a person the constraints change completely. See the AI pet video generator, which turns one photo of a dog or cat into a short clip with sound. For a still that should simply move rather than speak, bring photos to life is the lighter version of the same idea.
Which talking photo route to pick
You have a recording and the talking photo has to stay the same face. Avatar route. HeyGen at $29/mo if you also want voice cloning and long videos. TalkPix at $5 a pack if this is a one-off and a subscription is absurd for it.
You want a talking photo to say one line today and the voice does not matter. Video model. On Kavel that is Veo 3.1 Lite at 115 credits for 720p or 135 at 1080p, or Seedance 2.0 at 155 credits per second at 720p with audio in the same pass. New accounts start with 40 credits and the daily check-in adds 10 to 30 more, so the first clip needs no card. The pricing page has the credit packs, and the AI video maker is where you upload the still.
The photo is of a real person and you would rather not upload it anywhere. Local route: SadTalker or Wav2Lip with your own audio file.
On licences: Hedra lists commercial use on every paid tier, TalkPix grants it subject to your rights in the input, and DomoAI's paid tiers are watermark-free. Free tiers are where licences narrow, so read the tier you are on. And separately from any licence, you need the right to the face you upload. That is why the portrait here was generated rather than borrowed.

If the still should move without speaking at all, that is a different job. The image to video tools roundup covers it, and Kling 3.0 is the other model on our shelf that generates sound with the shot.
Talking photo FAQ
Is there a free app that makes a photo talk?
Several. HeyGen's free plan allows 3 videos a month at up to 1 minute, Vidnoz allows 5 photo avatars a day, and lip-sync.ai runs with no sign-up (all read that day). Free output is usually watermarked. On Kavel a new account starts with 40 credits and a daily check-in worth 10 to 30 more, which covers a first talking photo without a card.
Can AI take a picture and turn it into a video with voice?
Yes, in two different ways. An avatar tool builds the talking photo by syncing the picture to a voice you upload or choose. A video model such as Veo 3.1 or Seedance 2.0 generates the voice and the picture together in one pass, so you also get ambient sound — but no control over who it sounds like.
What AI turns a picture into a video for free without a watermark?
Among paid talking photo tiers, TalkPix says paid exports are clean and DomoAI advertises watermark-free outputs. Genuinely free and unwatermarked, the honest answer on 2026-09-11 is the local route, SadTalker or Wav2Lip on your own machine, or the sign-up credits a platform gives you, which are finite but do not stamp the output.
Audio, APIs and running it yourself
Can I upload my own audio and have a photo lip sync to it?
On the avatar route, yes: HeyGen accepts an uploaded audio file or a recording, Hedra takes an upload or a direct recording, and TalkPix accepts a script or audio. On the video-model route, no. Veo 3.1, Seedance 2.0 and Kling 3.0 generate the speech themselves and there is no field for your file.

Do these talking photo tools offer an API for creating videos programmatically?
HeyGen publishes a Talking Photo AI API on its pricing page and Hedra runs a developer platform alongside its studio, both read the same day. Video models are reached through their providers' APIs rather than a talking photo endpoint.
Can I run an open source talking head model locally, like SadTalker, LivePortrait or Wav2Lip?
Yes, with one correction worth knowing before you install anything. SadTalker and Wav2Lip are audio-driven: a still plus a recording gets you a talking photo. LivePortrait, despite appearing on every list beside them, is driven by a video rather than audio — its own repo describes it as portrait animation (all three repos read 2026-09-11). All are free per clip and upload the face nowhere, and all are meaningfully more work than opening a website.
Licences, phones and bad results
Can I use a talking photo video commercially?
It depends on the tier, not the tool. Hedra lists commercial use on all paid plans; TalkPix grants it subject to your rights in what you uploaded. Free tiers are where licences narrow.
Can I make a photo talk on my iPhone or Android, or does it need a desktop browser?
The avatar services and the video-model platforms in the table all run in a browser, so a phone works for the upload. The local route is the exception. SadTalker, Wav2Lip and LivePortrait are Python projects you install and run yourself, which is a desktop job.
Why does the mouth look wrong on my photo?
Usually the input. A talking photo needs a clean one, and a side profile, a shadow across the mouth, a low-resolution crop or two faces in frame are what vendor guidance tells you to avoid. If the input is clean and it still looks off, you are probably on a video model regenerating the face rather than animating it. That is the same mechanism that moved the eyebrows and the jawline in our own test above.
Sources for every talking photo figure on this page
Every figure here was read on 2026-09-11 from the page named: heygen.com/pricing, heygen.com/tool/create-ai-lip-sync-videos, heygen.com/avatars/avatar-iv, hedra.com/pricing, hedra.com/uses/ai-lip-sync, vidnoz.com/pricing.html, vidnoz.com/talking-head.html, talkpix.ai, domoai.app/pricing, lip-sync.ai, veed.io/pricing, and Kavel's own credit table and model configuration. The open-source rows were read from the SadTalker, LivePortrait and Wav2Lip repositories, and the three counter-examples from synthesia.io, d-id.com and captions.ai, all on 2026-09-11. The statements about which pages the answer engines cite, and the count of dollar figures and date markers on them, come from our own fetch of those eleven pages plus the ChatGPT and Perplexity fan-out recorded for this query on 2026-09-11. Second-hand figures are marked in the table. Our hands-on run was Veo 3.1 Lite, 1080p, image-to-video, 2026-09-11.





