How to Write Veo 3.1 Prompts for Product Ads
Quick answer: a Veo 3.1 prompt for a product ad is one paragraph written in a fixed order: camera behaviour, subject, physical action with the product, the spoken line in quotation marks, the audio bed, then the negatives. Veo maxes out at 8 seconds per clip(Google's API docs list 4, 6 and 8 seconds, and require 8 for 1080p or reference images), so an ad is built as a sequence of single-beat 8 second clips, not as one continuous take. Everything below is the prompt shape, the worked before-and-after examples, and the things Veo is genuinely bad at.
The six parts of a Veo 3.1 prompt, in order
Google's Veo documentation asks for subject, action, style, camera positioning, composition, focus and ambiance, and tells you to put speech in quotes. That is a general film-prompt framework. For a direct-response product ad it collapses into six slots. Fill every one of them, because anything you leave blank, Veo invents, and it usually invents a slow cinematic dolly move that makes the clip look like a car commercial rather than something filmed on a phone.
| Slot | What to write | What happens if you skip it |
|---|---|---|
| 1. Camera | Phone propped on a surface, locked frame, no pans or zooms | Veo adds drifting dolly and orbit moves. Instantly reads as stock footage. |
| 2. Subject and frame | Who, age range, where, how they are framed, what the light is | You get a studio-lit model in a grey void. |
| 3. Action with the product | One physical beat: hold it up, tear the seal, set it down | Hands float, the product changes shape mid-shot. |
| 4. Line in quotes | "The exact words, short enough to fit the clip" | The model improvises generic ad-speak, or mumbles. |
| 5. Audio bed | Voice quality, room tone, explicitly no music | Veo scores your UGC ad with a soundtrack. |
| 6. Negatives | No text, captions, subtitles, watermarks, fake UI | Burned-in captions you cannot remove without a re-render. |
Order matters more than you would expect. Veo conditions the performance on the prose that comes before the spoken words, so a delivery instruction placed after the quote is much weaker than the same instruction placed before it. That is why our own talking-actor prompt builder lifts delivery cues out of the script and inserts them as a sentence ahead of the words clause, rather than leaving them inline where the model can read them out loud.
Worked example one: the talking hook
Before (the prompt most people write):
A woman holding a skincare serum talking about how good it is.
That produces a well-lit stranger in a soft-focus room, holding a bottle that is not your bottle, saying words that are not your words, over library music.
After:
9:16 vertical phone-camera shot, one woman in her late twenties in a sunlit bathroom, framed chest-up, natural window light, no studio lighting. The phone is propped on the counter and the frame is locked: no pans, no zooms, no dolly. Slightly hurried, slightly amused delivery. She holds a frosted glass serum bottle up beside her cheek, label facing the lens, holds it still, and says to camera: "Three weeks. That is it. Three weeks and the texture on my cheeks is gone." Then she sets the bottle down on the counter. Audio: close conversational voice, faint bathroom room tone, no music. No on-screen text, no captions, no subtitles, no watermark, no stickers.
Every clause in the second version is doing a job you can point at. The framing sentence kills the cinematic camera. The delivery sentence sits before the quote so it shapes the read. The product action is one beat, held still, which is when Veo renders a label most reliably. The negative clause stops the burned-in captions.
Worked example two: the product beat with no face
Before:
Nice shot of the coffee bag on a counter.
After:
Locked overhead shot on a walnut kitchen counter, morning light from the left, no camera movement. Two hands enter the frame and slowly tear the gold foil seal of a matte black coffee bag; a few beans spill onto the wood. Shallow depth of field, the bag label stays in focus and stays still. Audio: the crisp tear of foil, beans rattling on wood, quiet kitchen room tone, no music, no voice. No on-screen text, no captions, no watermark.
Product-only beats are where Veo 3.1 is strongest, because there is no face and no speech to hold together. Native audio is doing real work here too: the foil tear and the beans are generated with the picture, so the sound is in sync without you cutting a foley track.
What the 8 second ceiling actually means for your ad
Google's Veo API documentation lists three durations, 4, 6 and 8 seconds, for both Veo 3.1 and Veo 3.1 Fast, and states that the 8 second setting is required when you use 1080p or reference images. Replicate's google/veo-3.1-fast page lists the same durations, 720p or 1080p, 16:9 or 9:16, 24fps, and up to three reference images.
So a 20 or 30 second ad is not one Veo prompt. It is three or four prompts, each carrying one beat, cut together. Write them as a shot list before you generate anything:
| Clip | Beat | Prompt focus |
|---|---|---|
| 1 (0 to 8s) | Hook, spoken to camera | Face, one line in quotes, product visible but not the subject |
| 2 (8 to 16s) | Product demonstration | No face, one physical action, native sound effects, no voice |
| 3 (16 to 24s) | Proof or objection handling | Same person, same room, same light, second line in quotes |
| 4 (24 to 30s) | Call to action | Face, product held up, short direct line, no music |
The hard part of multi-clip Veo is continuity. Reuse the same reference image, repeat the wardrobe, room and lighting description word for word across clips, and keep the same aspect ratio. Do not describe the room differently in clip three because you got bored of the sentence. If you want the whole 30 seconds spoken by one continuous person instead, Veo is the wrong tool and a longer-form talking model is the right one, which is the subject of our talking-head UGC ads guide.
What Veo 3.1 is bad at
An honest list, because the failure modes are predictable and you can design around them:
- Packaging text. Model-generated labels warp and mis-spell. If the brand name has to be legible, hold it large and static, or cut to a real photo for that beat.
- Long spoken lines. Eight seconds is roughly 15 to 20 spoken words. Push past that and the model rushes, drops clauses or invents its own ending.
- Fast hand work. Pouring, unboxing, applying and pressing all degrade quickly. One slow deliberate action per clip is the reliable pattern.
- Continuity between clips. There is no memory across generations. Same face, same room, same light only happens because you wrote it identically each time and reused the reference image.
- Uninvited on-screen text. It will sometimes add its own captions or a watermark-looking mark unless the negative clause is there.
We learned one thing the hard way that is worth passing on: appending your own framing text to a user-written prompt is not free. When we tested wrapping raw prompts in extra boilerplate before dispatch, the added text tripped model content filters more often than it helped, so our raw prompt mode now sends exactly what you typed, byte for byte, and the automatic framing only applies in talking-actor mode. If you are writing your own prompt, you own the negatives.
Which Veo setting should you pick?
For paid social: 8 seconds, 1080p, 9:16. Google's docs require 8 seconds for 1080p anyway, and 9:16 is the placement you are buying. Here is what each setting costs on our live catalog, so you can see why the short durations are a testing tool rather than a saving:
| Setting | Credits | Cost on Starter ($49 / 5,000 credits) | Use it for |
|---|---|---|---|
| Veo 3.1, 4s | 120 | ~$1.18 | Hook-only tests, no dialogue |
| Veo 3.1, 6s | 180 | ~$1.76 | One short line, one action |
| Veo 3.1, 8s (720p or 1080p) | 245 | ~$2.40 | Everything you actually ship |
Note that 720p and 1080p cost the same in credits on our side, because the underlying per-second rate is flat across resolutions. There is no reason to render a shippable ad at 720p.
A prompt template you can copy
Paste this, replace the bracketed parts, delete nothing:
9:16 vertical phone-camera shot, [one person, age range] in [room], framed chest-up, [light source], no studio lighting. Phone propped on a surface, frame locked, no pans, no zooms, no dolly. [Delivery: e.g. excited, fast-paced, slightly amused.] They [one physical action with the product, held still], and say to camera: "[15 to 20 words]". Audio: close conversational voice, quiet [room] ambience, no music. No on-screen text, no captions, no subtitles, no watermark, no stickers, no fake app interface.
If you want the same discipline applied to the words rather than the visuals, our 30 second UGC script guide covers the line-level structure, and Kling 3 vs Veo 3.1 covers when a different model is the better call for the same shot.
Veo 3.1 costs 245 credits for an 8 second render on UGC Vids AI, with the exact credit cost shown before you press generate. Free for 3 days. Cancel anytime.
Frequently asked questions
How do you write a good Veo 3.1 prompt for a product ad?
Write it as one paragraph in a fixed order: framing and camera behaviour, then the subject, then the physical action with the product, then the spoken line in quotation marks, then the audio bed, then what you do not want. Google's Veo documentation asks for subject, action, style, camera positioning, composition, focus and ambiance, and it specifically says to use quotes for speech. A prompt that names the camera behaviour and the product handling beats a prompt that only describes a mood, because Veo will invent camera moves and product handling if you leave them blank.
How long can a Veo 3.1 clip be?
Eight seconds. Google's Veo API documentation lists 4, 6 and 8 second durations for both Veo 3.1 and Veo 3.1 Fast, and notes that 8 seconds is required when you use 1080p or reference images. There is no setting that produces a 15 or 30 second Veo clip in one call, so a longer ad is built by generating several 8 second clips and cutting them together.
Can Veo 3.1 speak your exact script?
It generates the voice itself, natively, and Google's documentation tells you to put speech in quotation marks. In practice the model follows short lines closely and drifts on long ones, because it is composing the performance and the audio at the same time rather than reading a supplied voice track. Keep a spoken line to roughly the number of words a person says in the clip length you picked, around 15 to 20 words for an 8 second clip, and it will usually land verbatim.
Does Veo 3.1 add captions or subtitles you did not ask for?
It can. Video models sometimes burn in their own on-screen text, watermarks or fake app interface elements, and Veo is no exception. The fix is to state it in the prompt: no on-screen text, no captions, no subtitles, no watermark, no stickers, no fake social interface. In our talking-actor mode we append that negative clause automatically; in raw prompt mode you are responsible for it, because we send your prompt to the model verbatim.
What resolution and duration should I pick for a Veo 3.1 ad?
For a paid social ad, 8 seconds at 1080p in 9:16. Google's docs require the 8 second duration for 1080p output anyway, and 9:16 is the placement you are buying. On UGC Vids AI a Veo 3.1 generation costs 120 credits at 4 seconds, 180 at 6 seconds and 245 at 8 seconds, the same at 720p or 1080p, so the shorter durations only make sense when you are testing a hook idea and do not need the extra beat.
Why does my Veo 3.1 product shot get the product wrong?
Veo generates pixels rather than compositing your asset, so a product described only in words comes back as the model's idea of that product. Supply a reference image of the actual product, keep the label movement slow and deliberate in the prompt (held still beside the face, turned once toward the lens), and avoid asking for fast handling, pouring or unpacking in the same clip. Text on packaging is the most common failure, so frame the label large and steady, or cut to a real product photo for the beat that has to show the name.
Definitions
Compare alternatives
Stop reading. Start shipping.
Generate your first UGC ad in 2 minutes. No editing required.
Try the free generator