The split that decides it: verbatim or prompt-guided
Voice-driven models take a photo of a person and an audio track, then animate that person delivering the audio. The words in the output are the words in the track, because the model is not writing speech, it is performing speech you supplied. OmniHuman 1.5, VEED Fabric 1.0 and Pruna Avatar work this way.
Prompt-driven models take a photo and a text prompt and generate a scene, including their own audio. If your prompt describes a person saying something, they will generate a person saying something close to it. Kling 3.0, Happy Horse 1.1, Veo 3.1, Seedance 2.0, Seedance 1.5 Pro and Grok Video work this way.
For a talking-head ad, that difference is not a nuance. If your offer is 20% off and the model renders a person saying 25% off, you have an ad you cannot run. Every recommendation below starts from the voice-driven side for exactly that reason, and only moves to the prompt-driven side when the exact words genuinely do not matter.
Our pick: OmniHuman 1.5 for a standard 15 to 30 second read
OmniHuman 1.5 is the default answer for direct response talking heads. It renders up to 30 seconds in one continuous take, which covers a standard ad read without a cut, and it accepts any still image of a person as the subject.
It costs 490 credits for a 15 second read and 980 for 30, about 33 credits per spoken second. Its 720p and 1080p tiers are priced identically, so you take 1080p by default and the crop headroom comes for free. On the $49 Starter plan that is roughly $4.80 for a 15 second ad and $9.60 for a 30 second one, from a pool of 5,000 credits with no video count limit attached.
It is also the model we pin to a fixed build rather than letting it track whatever the vendor currently serves. If a spokesperson has to look the same in week six as in week one, that stability is doing real work.
What it is bad at: it animates a face from a photo and nothing else. The background is whatever was in your source image, there is no camera movement, there is no product being handled, and there is no scene. It also stops hard at 30 seconds. If your ad idea involves anything happening in the frame besides a person talking, this is the wrong model and no amount of prompting fixes that.
For scripts past 30 seconds: VEED Fabric 1.0
VEED Fabric 1.0 is the same shape of model, image plus voice track, built for longer runtimes. Its length follows the audio up to a 120 second ceiling, well past OmniHuman's, so a 45 second explainer or a 90 second founder story renders as one take rather than a stitch.
It costs 460 credits for a 15 second read and 920 for 30, about 31 credits per spoken second, so inside the range they share it is marginally cheaper than OmniHuman.
What it is bad at: resolution. Fabric renders 720p only, with no 1080p tier to buy. For a vertical feed placement at native size that rarely matters. For a file you plan to crop into several placements or punch into during the edit, it is a real ceiling, and it is the reason we do not make it the default pick for short reads.
The budget pick: Pruna Avatar at about 18 credits a spoken second
Pruna Avatar is the cheapest way to get a person saying your exact script. It runs its own text to speech, so it takes the script directly rather than a separately generated voice track, and it prices at about 18 credits per spoken second. A 15 second read lands near 275 credits against OmniHuman's 490, so you get close to two renders for the price of one.
It is sold on a single 1080p tier, and because it has no duration input the output length follows the script exactly like the other two.
What it is bad at: voice range and length. Because the speech is generated inside the model, you get a preset male or female voice in your selected language rather than a wide voice library, so it is the weakest of the three if a specific vocal character is part of your brand. Its speech also truncates around 60 seconds, so it is not the model for long form. And like the rest of this tier, it animates a portrait and cannot stage a scene.
Where it earns its place is volume. If you are testing eight different script angles for the same offer before deciding which one to put spend behind, paying 18 credits a second instead of 33 for the test renders is a straightforwardly better use of the pool.
The shortlist, side by side
Three models can recite your script. Here is the whole decision in one table.
When a prompt-driven model is actually the better call
There is a version of a talking-head ad where the words are not load-bearing: a person on camera reacting, gesturing, holding the product, moving through a room, with the message carried by captions and the edit rather than by the audio. That ad wants a scene, and the voice-driven tier cannot build one.
For that, Happy Horse 1.1 renders up to 15 seconds with native audio and accepts up to nine reference images, at 140 to 420 credits at 720p or 180 to 550 at 1080p. Kling 3.0 renders up to 15 seconds with a 4K tier and seven reference slots, at 260 to 780 credits in 1080p. Veo 3.1 Fast renders 4, 6 or 8 seconds at 120, 180 and 245 credits with 1080p included, which suits a short spoken beat as long as you accept that the beat is 8 seconds and the wording is approximate.
The honest framing is that these are not competing for the same job. They are the right answer to a different brief that happens to also have a person in it.
Models we would not use for this job
- Kling 2.6: our picker files it under b-roll and scenes rather than talking actors, and its dialogue would be prompt-guided in any case. Good for product motion, wrong for a read.
- Grok Video: caps at 720p and 10 seconds and its dialogue is prompt-guided. Excellent for cheap concept sweeps, not for a scripted read.
- Seedance 1.5 Pro: 12 second ceiling with prompt-guided dialogue at 20 credits a second. The cheapest 1080p second among the prompt-driven models, but not a script-accurate one.
- Sora 2 and Sora 2 Pro: OpenAI is retiring both on 24 September 2026, so we would not build a talking-head workflow on them now regardless of how they perform. Their absence from this ranking is deliberate.
How to decide in one pass
Count the words in your script. At roughly 2.3 spoken words a second, under about 69 words is a 30 second read or shorter, and that is OmniHuman 1.5 territory. Over that and you want VEED Fabric 1.0.
Then ask whether this render is a test or a deliverable. Tests belong on Pruna Avatar at about 18 credits a second, where you can afford to try six script angles. Deliverables belong on OmniHuman 1.5 at 1080p.
Finally, check whether anything in the frame moves besides the person. If it does, you are not making a talking-head ad and you should be looking at Kling 3.0 or Happy Horse 1.1 instead.
All of these models share one credit balance on UGC Vids AI, priced at $49 for 5,000 credits, $99 for 12,000 or $199 for 25,000 with 30% off annually, and no cap on how many videos those credits become. A single Starter month covers hearing your own script on all three. The trial is $1 for 7 days. Cancel anytime.