Neither model has a duration setting
OmniHuman 1.5 and VEED Fabric 1.0 both work the same way at the input level: one image of a person, one audio track. There is no seconds parameter to set because the audio already contains the answer. Feed in a 22 second voice track and you get a 22 second video.
Because of that, billing follows the script rather than a tier. UGC Vids AI estimates the spoken length of your script at roughly 2.3 words a second and charges per spoken second, so a 40 word script prices as about 17 seconds and a 70 word script as about 30.
The practical consequence is that writing tighter copy is a direct cost saving on these two models in a way it is not on a fixed-tier model. Cutting fifteen words off a script saves you real credits, not just runtime.
OmniHuman 1.5: up to 30 seconds, with a 1080p tier
OmniHuman 1.5 is ByteDance's avatar animation model. On UGC Vids AI it renders up to 30 seconds in one continuous take, which is its hard ceiling. Push a longer script at it and you are past what the model accepts.
Pricing is 490 credits for a 15 second read and 980 for 30, which is about 33 credits per spoken second, or roughly $4.80 and $9.60 on the $49 Starter plan. The detail worth knowing is that its 720p and 1080p tiers cost exactly the same number of credits, so there is no reason to ever choose 720p on this model. Effectively you are buying a 1080p talking head.
It is also the one model in the lineup pinned to a specific build rather than tracking whatever the vendor currently serves as default. That is a stability decision on our side: the version you rendered with last month is the version you render with today, which matters when a spokesperson has to look consistent across a campaign that runs for weeks.
VEED Fabric 1.0: no 30 second wall, 720p only
VEED Fabric 1.0 is the long-form option. It is built for extended talking video, and its length follows the voice track up to a ceiling of 120 seconds, which at about 2.3 spoken words a second is a script of roughly 276 words. If yours is a 45 second explainer or a 90 second founder story, this is the model that renders it in one piece.
Its cost is 460 credits for a 15 second read and 920 for 30, about 31 credits per spoken second, so it is marginally cheaper than OmniHuman across the range they share.
The trade is resolution: Fabric is sold at 720p only. There is no 1080p tier to buy. For a feed placement that is often irrelevant, but if the file needs to be cropped, punched into, or delivered somewhere demanding more pixels, that ceiling is the thing that will bite.
Script length decides it
At roughly 2.3 spoken words a second, here is where the answer lands.
The 1080p question, and when it is worth 30 credits
Across the 15 second read that most direct response ads land on, OmniHuman costs 490 credits and Fabric costs 460. That is a 30 credit gap, about 29 cents on the Starter plan, and it is the entire price of moving from a 720p ceiling to a 1080p one.
If the finished ad goes straight into a vertical feed placement at the size it was rendered, the gap buys you very little. If you crop the same file for a square placement, add captions with a safe margin, or punch in on the speaker's face for the second half, it buys you headroom you cannot get back later.
That is genuinely the whole financial comparison between these two inside the 30 second window. They are within about 6% of each other per spoken second, so this is not a budget decision. It is a resolution decision with a rounding error attached.
The voice is not part of this choice
A common assumption is that switching between these two models changes how the ad sounds. It does not. Both are driven by a voice track that is generated before the render, from your script and the voice you selected, and then handed to whichever model you picked.
That means the voice, the language and the delivery are settled at the voice step rather than the model step, and swapping models leaves the audio identical. If you rendered a read on OmniHuman and want to see it at a longer length on Fabric, the person sounds the same.
It also means that on both models the words are exactly the words you wrote. That is the defining property of the voice-driven tier and the reason it exists: prompt-driven models such as Kling 3.0, Veo 3.1 or Happy Horse 1.1 generate dialogue guided by your prompt, which is fine for atmosphere and unacceptable for a claim your compliance team signed off on.
Which model for which ad job
- A standard 15 to 30 second direct response read: OmniHuman 1.5, for the 1080p tier at effectively the same price.
- An explainer, a founder story or a walkthrough past 30 seconds: VEED Fabric 1.0, which is the only one of the two that gets there in one render.
- A spokesperson who must look identical across a multi-week campaign: OmniHuman 1.5, since it is pinned to a fixed model build.
- Long-form organic content repurposed into ads: VEED Fabric 1.0, where the 720p ceiling matters less and the length ceiling matters more.
- Anything that will be re-cropped into several placements: OmniHuman 1.5 at 1080p, for the crop headroom.
- Product motion, camera moves or scenes: neither. Both animate a face from a still photo, so a demo of an object in motion belongs on a prompt-driven model.
Verdict: count the words first
Under about 69 words, use OmniHuman 1.5. The price difference against Fabric is small enough to ignore and the 1080p tier is free relative to its own 720p tier, so you are taking resolution at no real cost.
Over about 69 words, the question answers itself, because OmniHuman stops at 30 seconds and VEED Fabric 1.0 does not. There is no clever workaround here that beats simply using the model built for the length.
Both sit in the same generator on the same credit balance at UGC Vids AI, so nothing forces you to pick in advance. A Starter month covers hearing the same script at both lengths on your own avatar. The trial is $1 for 7 days. Cancel anytime.