A useful AI video prompt names the subject, action, camera behavior, setting and visual style, plus audio when the selected model supports it. For image-to-video, focus on what should change in the uploaded image rather than describing what is already visible. Clear motion instructions usually matter more than adding more adjectives.
Compare these two prompts.
Before
A cinematic luxury perfume ad.
After
A clear glass perfume bottle stands on black stone. Fine mist drifts behind it while a narrow warm light moves across the glass. Slow 90-degree camera orbit, macro commercial photography, deep black background, restrained reflections. No generated text.
The second prompt gives the model a shot to execute rather than a mood to guess.
What should an AI video prompt include?
A strong AI video prompt should define five things: subject, action, camera, setting and style, and audio when supported.
Use this structure:
- Subject: Who or what is in the shot?
- Action: What changes during the video?
- Camera: How is the shot framed and how does the camera move?
- Setting + style: Where does it happen and what should it look like?
- Audio: What should the viewer hear?
Five-field prompt template
Subject: [main person, object, or product]
Action: [one clear movement or sequence]
Camera: [shot size + camera movement]
Setting + style: [location, lighting, palette, texture, mood]
Audio: [dialogue, ambience, sound effects, music, or silence]
Filled example
Subject: A matte cobalt-blue water bottle on a pale stone plinth.
Action: Condensation slowly forms and catches the light while the bottle remains still.
Camera: Medium product shot with a slow 120-degree clockwise orbit.
Setting + style: Minimal studio, soft side light, cool grey background, premium commercial photography.
Audio: Quiet studio ambience and one soft glass-like tonal hit at the end.
Do not add a field just because the template contains it. If sound is irrelevant, leave it out. If an uploaded image already establishes the product, spend more of the prompt on motion.
How do you write a text-to-video prompt?
Start with one shot objective, define its motion, then add camera and visual direction only where those choices matter.
A practical workflow is:
- Decide what must happen.
- Name the subject.
- Add one main action.
- Define framing and camera movement.
- Add lighting and environment.
- Add sound only when supported and necessary.
- Generate, inspect the failure, then revise only that part.
Example:
A woman in a red raincoat walks through a quiet station at night.
Medium tracking shot moving beside her at walking speed.
Wet floor reflections, cool fluorescent lights, restrained cinematic contrast.
She looks toward an arriving train near the end of the shot.
You can test this structure in the AI video generator.
Why should image-to-video prompts focus on motion?
The uploaded image already defines appearance, so tell the model what should move and what must remain unchanged.
Instead of:
A blue bottle with a white logo on a stone background.
Use:
Keep the bottle design, logo, color, proportions and background unchanged.
Add a gentle camera push-in while a narrow highlight moves slowly from left to right across the bottle.
No shape change and no generated lettering.
For more complex reference-driven movement, see AKOOL's Video Reference guide.
What should you put in negative instructions?
Use negative instructions to protect a few likely failure points:
No generated text.
No logo deformation.
No extra objects entering frame.
No camera shake.
No shape morphing.
Do not write a second prompt made entirely of negatives.
AI video prompts for product ads
Product hero shot prompt
1. Cobalt bottle orbit
Use case: Premium product reveal
Input: Product image
Recommended model: Seedance 2.0
Aspect ratio: 16:9
A matte cobalt-blue water bottle stands on a pale stone plinth. Fine condensation beads catch soft side light. The camera makes a slow 120-degree clockwise orbit while the bottle remains centered and completely still. Minimal grey studio, realistic materials, controlled reflections. Keep the label shape unchanged. No text overlay.
Expected result: A polished studio product shot with controlled parallax and readable product geometry. The most likely weakness is logo or silhouette drift toward the end of the orbit.
Adjustment: Reduce the orbit to 60 degrees if geometry changes.
2. Skincare macro detail
Use case: Packaging texture
Input: Product photo
Recommended model: Seedance 2.0
Aspect: 1:1
Keep the serum bottle exactly as shown. Macro push-in toward the glass dropper while tiny condensation droplets form on the surface. Soft warm edge light, shallow depth of field, clean cream background. No label changes, no additional objects.
Expected result: Strong packaging preservation with subtle surface motion. Condensation may interfere with fine label text.
Adjustment: Remove condensation if text stability matters more than atmosphere.
3. Running shoe lifestyle shot
Use case: Sports ad
Input: Product reference
Recommended model: Kling 3.0
Aspect: 16:9
The same running shoe lands on wet pavement at dusk. Low-angle tracking shot moves beside the athlete's foot. One realistic splash on impact, then the runner pushes forward. Warm storefront reflections against cool blue street light. Preserve the shoe design and color.
Expected result: Energetic lateral movement and a strong impact beat. The sole and laces are the highest-risk details.
Adjustment: Limit the action to one foot strike if shoe geometry drifts.
4. Cosmetic lighting transformation
Use case: Before/after visual without a performance claim
Input: Text or references
Recommended model: Veo 3.1
Aspect: 9:16
Begin with a plain neutral vanity setup. A cosmetic jar slides into frame. Lighting gradually changes from flat daylight to polished beauty lighting while the same objects remain in place. End on a clean hero composition. No skin transformation claims.
Expected result: A clear aesthetic transformation driven by light rather than an unsupported product claim.
Adjustment: Keep the jar static and change only lighting if product geometry shifts.
5. Vertical sneaker hook
Use case: Reels/TikTok
Input: Product image
Recommended model: Kling 3.0
Aspect: 9:16
9:16 close-up of a sneaker landing on wet pavement at dusk. Quick low-angle tracking move, one splash in slow motion, warm storefront reflections. End on a still product frame for one second. Keep the shoe unchanged. No generated lettering.
Expected result: Immediate visual hook with a usable CTA hold.
Adjustment: Remove slow motion if the splash or shoe motion becomes inconsistent.
6. Unboxing detail
Use case: Packaging reveal
Input: Product and box references
Recommended model: Kling 3.0
Aspect: 9:16
Hands slowly lift the lid of a matte black product box. The product remains centered inside. Overhead close-up, soft diffused studio light, deliberate movement, no fast cuts. Keep all packaging geometry unchanged.
Expected result: Controlled overhead reveal. Fine finger contact is the likely failure point.
Adjustment: Replace human hands with a clean automatic lid lift if necessary.
7. Beverage pour
Use case: Food and beverage
Input: Product reference
Recommended model: Veo 3.1
Aspect: 16:9
A chilled glass bottle pours sparkling citrus drink into a clear glass filled with ice. Tight side close-up. Realistic bubbles and condensation, warm afternoon sunlight, slow camera push-in. Preserve the bottle label. No text overlay.
Expected result: Convincing fluid motion and rich commercial lighting. Rotating label details may soften.
Adjustment: Keep the bottle facing camera and shorten the pour.
8. Headphone turntable
Use case: Electronics
Input: Product image
Recommended model: Seedance 2.0
Aspect: 1:1
Keep the headphones exactly as supplied. The product rotates slowly 45 degrees on a dark studio platform while a cool rim light moves across the ear cups. Static camera, clean black background, precise commercial product lighting.
Expected result: Clean ecommerce-style product motion.
Adjustment: Move only the light instead of rotating the headphones if their structure changes.
9. Fashion accessory close-up
Use case: Luxury fashion
Input: Product image
Recommended model: Veo 3.1
Aspect: 9:16
A leather handbag rests on a brushed metal table. Slow dolly in as a narrow beam of afternoon light travels across the leather texture. Minimal gallery setting, soft shadows, no hands, no new objects. Preserve stitching, hardware, shape, and logo.
Expected result: Stable product with subtle premium movement.
Adjustment: Lock the camera and animate light only if metal hardware deforms.
10. Food hero shot
Use case: Restaurant ad
Input: Image-to-video
Recommended model: Veo 3.1
Aspect: 16:9
Use the uploaded burger photo as the exact food reference. Keep ingredient order and shape fixed. Add gentle steam rising from the top and a slow five-percent camera push-in. Warm restaurant lighting, dark background, no ingredient movement.
Expected result: Strong preservation because most motion occurs in steam and camera movement.
Adjustment: Remove steam if it distorts the food silhouette.
15-second product story prompt
11. Three-shot coffee story
Use case: Short narrative
Input: Product and character references
Recommended model: Kling 3.0
Aspect: 16:9
Create a short three-shot coffee story.
Shot 1: Wide morning kitchen. The same woman places a black ceramic mug on the counter.
Shot 2: Medium close-up as she pours coffee into the same mug.
Shot 3: Close-up of the mug and her hands as she lifts it.
Keep the woman, wardrobe, mug, kitchen and warm amber grade consistent.
Expected result: Coherent short sequence with stronger continuity than three unrelated prompts.
Adjustment: Generate the shots individually if identity or mug continuity drifts.
12. Tech product reveal with CTA hold
Use case: Paid social
Input: Product reference
Recommended model: Seedance 2.0
Aspect: 9:16
Start on an extreme close-up of the product texture. Pull back slowly to reveal the complete device on a clean desk. A hand enters once and presses the main control. End with the product alone and hold the final composition for two seconds. Preserve all design details. No generated text.
Expected result: Strong reveal and useful CTA ending.
Adjustment: Remove the hand interaction if contact with the control becomes unstable.
AI video prompts for social media
13. One-second visual hook
9:16. Open immediately on a red umbrella snapping open in heavy rain directly toward camera. Quick push-in, then hold for a beat. High contrast city reflections. No text.
Recommended model: Veo 3.1
Expected result: Fast, readable opening action.
Adjustment: Reduce rain density if it hides the umbrella.
14. Seamless coffee loop
Top-down coffee cup. Cream forms one clean spiral, reaches the exact starting pattern, and loops smoothly. Static overhead camera, warm café light.
Recommended model: Seedance 2.0
Expected result: Visually satisfying loop, though exact frame-perfect looping may need editing.
Adjustment: Reduce to one simple spiral.
15. Outfit transition
9:16 full-body fashion shot. The subject steps behind a narrow pillar wearing outfit A and emerges from the other side wearing outfit B. Fixed camera and identical walking speed.
Recommended model: Kling 3.0
Expected result: Clean reveal with some risk of facial or garment drift.
Adjustment: Use first/end references where available.
16. Recipe hook
Extreme close-up of hot chili oil hitting freshly cooked noodles. One fast pour, visible steam, realistic food texture, quick handheld push-in.
Recommended model: Veo 3.1
Expected result: Strong texture and fluid motion.
17. Desk makeover
Locked 9:16 camera on a messy desk. Objects slide smoothly into organized positions one group at a time. End on a clean workspace.
Recommended model: Kling 3.0
Expected result: Clear transformation. Object count may drift during complex movement.
18. Travel reveal
Start behind a stone doorway in shadow. The camera walks forward through the doorway to reveal a bright coastal landscape. Smooth forward movement, natural exposure transition, no cut.
Recommended model: Veo 3.1
Expected result: Strong reveal and exposure change.
19. Fitness motion
Side medium shot of an athlete performing one controlled kettlebell swing. Fixed camera, realistic body mechanics, gym ambience, no slow motion.
Recommended model: Veo 3.1
Expected result: Better consistency than a multi-repetition exercise prompt.
20. Creator reaction clip
9:16 close-up of a creator looking down at a laptop, pausing, then looking directly into camera with a surprised smile. Subtle handheld movement, daylight home office.
Recommended model: Kling 3.0
Expected result: Natural micro-performance with minor facial risk during the expression change.
21. Beauty transition
Close-up portrait. The subject turns her head slowly from left profile to camera while the lighting changes from cool window light to warm beauty light. Face and hairstyle remain unchanged.
Recommended model: Seedance 2.0 with reference image
Expected result: Stable identity if motion stays restrained.
22. Mini product demonstration
9:16 overhead shot. A hand places a wireless charger on a desk, then places a phone directly onto it. One clear interaction. Clean commercial lighting.
Recommended model: Veo 3.1
Expected result: Simple interaction, but hand-to-object contact requires close QA.
23. Street-style loop
A model walks past the camera from left to right. As the model exits frame, another identical composition begins from the left to create a clean loop. Neon evening street, steady camera.
Recommended model: Kling 3.0
Expected result: Usable looping source that may still require frame trimming.
24. Slow reveal hook
Begin completely out of focus on a brightly colored object. Rack focus over two seconds to reveal the product sharply, then hold. No camera movement.
Recommended model: Veo 3.1
Expected result: Controlled focus transition and simple product reveal.
Cinematic and camera-motion prompts
For deeper model-specific direction, see AKOOL's Kling 3.0 prompt guide.
25. Orbit
Medium shot of a ceramic sculpture. The camera makes a smooth 90-degree clockwise orbit while maintaining equal distance from the subject.
Expected result: Clear parallax and controlled circular movement.
Adjustment: Reduce to 45 to 60 degrees if geometry changes.
26. Dolly in
Wide shot of a lone diner in an empty late-night restaurant. Slow straight dolly toward the subject. No zoom.
Expected result: Gradual emotional emphasis without changing lens perspective.
27. Dolly out
Start on a tight portrait. Slowly dolly backward to reveal the character standing alone in a huge warehouse.
Expected result: Effective scale reveal. Background geometry is the main QA point.
28. Crane down
Begin high above a quiet courtyard. Slowly crane downward and forward until a single person becomes the visual focus.
Expected result: Strong establishing-to-subject transition.
29. Overhead tracking
Direct overhead shot following a cyclist along a straight painted lane. Camera tracks at constant speed directly above the cyclist.
Expected result: Stable graphic composition if the road remains simple.
30. Controlled handheld
Medium close-up of a chef closing a restaurant at night. Natural handheld micro-movement only. Follow the chef two steps toward the door.
Expected result: Natural documentary feel without excessive shake.
31. Low-angle tracking
Low camera beside a skateboard. Track parallel with the board as it rolls across smooth concrete. Keep the wheels and board geometry stable.
Expected result: Dynamic movement, with wheels as the main deformation risk.
32. Slow pan
Locked tripod position. Slow left-to-right pan across a quiet vintage record store, ending on one customer browsing a shelf.
Expected result: Predictable horizontal reveal.
33. Tilt reveal
Begin on a person's shoes. Slowly tilt upward to reveal the full outfit and face. Subject remains still.
Expected result: Reliable fashion reveal because subject motion is minimal.
34. POV walk
First-person camera walks slowly through a narrow greenhouse aisle. Plants pass naturally at both edges of frame.
Expected result: Immersive movement with possible edge morphing in foreground leaves.
35. Whip-pan transition
Begin on a daytime city street. Fast whip pan to the right creates motion blur, resolving into the same framing at night.
Expected result: Strong transition concept, but matching architecture on both sides is difficult.
Adjustment: Use start/end references where supported.
36. Static tension shot
Completely locked camera on an empty hallway. Nothing moves for three seconds. Then one door at the far end slowly opens.
Expected result: High consistency because almost the entire frame stays fixed.
Image-to-video and reference-driven prompts
37. Product photo push-in
Use the uploaded product photo as the exact appearance reference. Keep logo, color, proportions and silhouette fixed. Add a gentle camera push-in and moving highlights.
Recommended model: Seedance 2.0
Expected result: Strong product preservation with restrained motion.
38. Portrait blink
Preserve the uploaded person's face, hairstyle, clothing and background. Add one natural blink and a slight head turn toward camera.
Recommended model: Seedance 2.0
Expected result: High identity consistency due to minimal movement.
39. Packaging rotation
Keep all packaging text and colors unchanged. Rotate the box slowly 20 degrees clockwise, then stop.
Expected result: Usable subtle rotation. Side-panel text may become less stable.
40. Fashion still to moving portrait
Preserve the model's identity and outfit. Add a small shift of body weight and one natural fabric movement from a light breeze.
Expected result: Natural portrait animation with low continuity risk.
41. Food photo animation
Keep the exact plate, food arrangement and background. Add gentle steam and a subtle camera push-in.
Expected result: Strong source-image preservation.
42. Architecture photo
Preserve the building geometry. Add slow cloud motion and subtle movement in nearby trees. Static camera.
Expected result: Stable architecture because motion is confined to environment.
43. Mascot consistency
Use the uploaded mascot as the identity reference. Keep face, proportions, colors and costume fixed. The mascot waves once and takes one step toward camera.
Expected result: Recognizable mascot identity with moderate limb-motion risk.
44. Product start-to-end frame
Begin exactly from the supplied first frame and end on the supplied final frame. Move gradually from the wide product composition to the close hero composition.
Expected result: More controlled composition transition when the selected model supports first/end frames.
45. Motion reference transfer
Use the uploaded product or character reference for appearance. Follow the pacing and camera movement of the supplied reference clip while preserving the new subject's identity and design.
Expected result: Stronger motion control than text alone, with some risk that subject proportions follow the motion reference too closely.
See the Video Reference guide for setup detail.
46. Multi-angle product continuity
Use the supplied product reference images to preserve shape, materials and branding. Create a slow three-quarter tracking shot around the product.
Recommended models to compare: Seedance 2.0, Kling 3.0, Veo 3.1
Expected result: Better product identity than text-only generation, but unseen surfaces may still be reconstructed differently.
Dialogue and sound prompts
47. Short dialogue
Medium close-up of a café owner standing behind the counter at closing time. She looks into camera and says, "See you tomorrow." Natural speaking pace. Quiet café room tone and a distant chair scrape. No music.
Recommended model: Veo 3.1
Expected result: Short line should fit comfortably, with room for natural ambience.
48. Ambient product sound
Close-up of sparkling water pouring into a glass with ice. Slow camera push-in. Audio: clear liquid pour, light ice movement, faint outdoor ambience. No dialogue and no music.
Recommended model: Veo 3.1
Expected result: Audio should reinforce visible actions without competing elements.
49. Beat-synced fashion clip
9:16 fashion clip synchronized to the uploaded beat. Hold the first pose until the first strong beat, then cut to a low-angle walking shot. End on a still close-up on the final beat.
Recommended workflow: Seedance workflow with audio reference if exposed in the current UI.
Expected result: Broad beat alignment rather than frame-perfect synchronization.
Adjustment: Perform final timing in an editor if exact beat cuts are required.
50. Intentional silence
Wide static shot of an empty snowy road at dawn. One person walks slowly across frame in the distance. No music, no dialogue, no artificial sound effects. Preserve quiet natural ambience only.
Recommended model: Native-audio model with ambient-sound support
Expected result: Sparse environmental audio with no unnecessary soundtrack.
Downloadable AI video prompt structure
SUBJECT
Who or what is the shot about?
ACTION
What changes during the clip?
CAMERA
Shot size:
Camera position:
Camera movement:
Movement speed:
SETTING + STYLE
Location:
Time of day:
Lighting:
Color palette:
Visual treatment:
AUDIO
Dialogue:
Ambience:
Sound effects:
Music:
Silence:
CONTINUITY
What must stay unchanged?
NEGATIVE INSTRUCTIONS
Only list likely failure modes.
For image-to-video, shorten the subject section and expand Action, Camera, and Continuity.
Why did my AI video prompt fail?
| Problem | Weak instruction | Better instruction | Better workflow |
|---|---|---|---|
| Frozen motion | “Cinematic woman in a café” | “She turns toward the window, then lifts the cup once” | Choose a stronger motion model |
| Subject drift | “Same product throughout” | “Use uploaded product as reference; preserve logo, color and silhouette” | Reference-driven generation |
| Unwanted text | “Luxury ad with tagline” | “Leave negative space. No generated lettering” | Add real copy in post |
| Camera ignored | “Dynamic camera” | “Slow clockwise 60-degree orbit at constant distance” | Video Reference |
| Too much action | Five actions in six seconds | One primary action per shot | Split the sequence |
| Audio mismatch | Long dialogue + action | One short attributed line | Separate audio if needed |
| Product changes | Full rotation + reflections | Smaller camera move | More references |
| Unnatural hands | Fine motor task | Simplify contact | Insert or real footage |
Before and after: frozen motion
Before
A man in a stylish kitchen, cinematic.
After
Medium shot of a man standing at a kitchen counter. He picks up one red apple, turns it once in his hand, then places it beside a cutting board. Camera remains fixed. Soft morning window light.
Before and after: unwanted product changes
Before
Make this bottle look dramatic while the camera moves around it.
After
Use the uploaded bottle as the exact appearance reference. Keep the silhouette, cap, label and color unchanged. Move the camera slowly 45 degrees clockwise while maintaining the same distance. No product deformation and no generated text.
If prompt revisions keep failing, change the workflow rather than making the prompt longer.

