Veo 3.1 JSON Prompts: Templates and Examples for AI Video

Updated: 
October 5, 2026
Learn a clear JSON structure for Veo 3.1 video prompts, with examples for shots, motion and audio. Convert each template into an AKOOL-ready prompt.
Table of Contents

A Veo JSON prompt organizes a shot into fields such as subject, action, camera, lighting, and audio. The model interface determines whether you can paste JSON directly. In AKOOL, use JSON as a planning format and translate its values into a normal video prompt unless the current interface explicitly documents structured JSON input.

If you searched for a veo3 json prompt generator, this distinction matters. JSON can make a creative brief easier to read, reuse, and validate, but it is not automatically a special Veo prompting language.

What is a Veo 3.1 JSON prompt?

A Veo 3.1 JSON prompt is best understood as a structured way to organize creative instructions unless the interface or API you are using explicitly defines those JSON fields.

For example:

{
  "subject": "blue water bottle on a stone table",
  "action": "condensation rolls slowly down the bottle",
  "camera": "slow close-up push-in",
  "lighting": "soft morning window light",
  "audio": "quiet room tone, no dialogue"
}

This is valid JSON. It is also easy for a human or another system to read.

But it is not an official Veo schema just because the keys are named subject, camera, or lighting.

Google's actual Gemini API uses a structured API request with documented fields such as prompt, aspectRatio, durationSeconds, and resolution. Those parameters belong to Google's API contract, not to a universal JSON prompting syntax.

The planning JSON above can be converted into this ordinary prompt:

A blue water bottle sits on a stone table. Condensation rolls slowly down the bottle while the camera makes a gentle close-up push-in. Soft morning window light falls across the scene. Audio is quiet room tone with no dialogue.

That is the core workflow of json prompting veo3: structure the idea first, then send the model the form of input the actual interface accepts.

Unsupported keys may simply have no effect if the application does not parse them.

How should you structure a Veo video prompt?

A useful planning schema separates subject and action from camera, lighting, and sound. Keep interface controls such as duration, resolution, and aspect ratio outside the creative JSON unless the specific product documents those fields.

Subject and action

The subject tells Veo what is present. The action tells it what changes during the clip.

{
  "subject": "chef standing alone in a small restaurant",
  "action": "the chef turns off the final pendant light and walks toward the back door"
}

Use concrete movement. “Looks cinematic” gives the model less temporal information than “turns off the light and walks toward the door.”

Camera and lighting

Camera fields are useful because they force you to distinguish subject motion from camera motion.

{
  "camera": "slow dolly backward at eye level",
  "lighting": "warm tungsten light fading into cool blue alley light"
}

Avoid vague values such as "camera": "dynamic".

Use physical instructions such as push-in, pan, orbit, tracking shot, crane, or static frame.

Dialogue and sound

Veo 3.1 supports native audio. Google's current documentation says Veo 3.1 generates audio with video, including prompt-directed dialogue and sound.

Keep dialogue attributed:

{
  "dialogue": [
    {
      "speaker": "friend A",
      "line": "We made it."
    }
  ],
  "ambient_audio": "soft cafe chatter and a distant cup clink"
}

Do not mix three speakers, music, complex Foley, and multiple simultaneous actions into a short clip unless you have a reason to test that complexity.

Aspect ratio and duration

Treat duration, aspect ratio, and resolution as interface or API settings, not as assumed creative JSON keys.

Google's current API documentation supports 4, 6, or 8-second Veo 3.1 generation, 16:9 and 9:16 aspect ratios, and 720p, 1080p, or 4K depending on configuration.

AKOOL's current public generator page lists Veo 3.1 at up to 8 seconds, 16:9 and 9:16, with native audio and up to 4K.

Those are platform settings. Do not assume this planning object controls them:

{
  "aspect_ratio": "9:16",
  "duration": "8s",
  "resolution": "4K"
}

unless the current AKOOL UI or API explicitly documents those keys.

How do you use JSON-style prompts in AKOOL?

Use JSON to plan the shot, then convert the values into a concise natural-language brief before generating in AKOOL. AKOOL's public documentation currently confirms Veo 3.1 availability, but it does not document raw JSON as a dedicated prompt-input mode.

A practical mapping looks like this:

JSON fieldPlanning valueAKOOL-ready prose
subjectblue bottle on stone table“A blue bottle sits on a stone table”
actioncondensation rolls downward“Condensation slowly rolls down the bottle”
cameraslow close-up push-in“The camera slowly pushes closer”
lightingsoft morning window light“Soft morning window light falls from the left”
audioquiet room tone“Quiet room tone, no dialogue”

Combined:

A blue bottle sits on a stone table. Condensation slowly rolls down the bottle while the camera makes a gentle close-up push-in. Soft morning window light falls from the left. Quiet room tone, no dialogue.

To generate a Veo video in AKOOL, choose Veo 3.1 from the current model selector, enter the converted prompt, then configure the options exposed by the interface. AKOOL currently positions Veo 3.1 as a model for lifelike motion and native sound.

For broader direction on camera language, references, and cinematic prompting, use AKOOL's general Veo 3.1 prompt guide. This article stays focused specifically on JSON organization and conversion.

Required screenshot before publication: current AKOOL generator with Veo 3.1 selected, prompt field visible, and the available duration, aspect-ratio, quality, and audio controls visible.

What are five copyable Veo 3.1 JSON prompt examples?

The following five objects are parser-valid JSON. The keys are an editorial planning schema, not a claim that AKOOL or Veo officially requires these exact field names.

1. Product ad

Use case: premium product shot

{
  "subject": "matte cobalt-blue water bottle on a pale stone table",
  "action": "condensation rolls slowly down the bottle while the bottle remains still",
  "camera": "slow close-up push-in",
  "lighting": "soft morning window light from camera left",
  "audio": "quiet room tone, no dialogue"
}

Natural-language version:

A matte cobalt-blue water bottle sits on a pale stone table. Condensation rolls slowly down the bottle while it remains still. The camera makes a slow close-up push-in. Soft morning window light comes from camera left. Quiet room tone, no dialogue.

AKOOL settings: Select Veo 3.1 and set duration, aspect ratio, and quality using the current UI.

Tested output: The generated clip should keep the water bottle stable while condensation moves naturally and the camera performs a slow push-in. Check whether the bottle's color, shape, position, and condensation remain consistent throughout the shot.

Revision to test:

Keep the bottle completely static. Reduce the camera movement to a very slow 10 percent push-in. Preserve the bottle color and silhouette.

Media: Insert generated clip plus representative frame.

2. Talking scene

Use case: short dialogue

{
  "scene": "two friends sitting across from each other in a quiet cafe",
  "shot": "medium two-shot, static camera",
  "action": "friend A looks up from the table and smiles",
  "dialogue": [
    {
      "speaker": "friend A",
      "line": "We made it."
    }
  ],
  "ambient_audio": "soft cafe chatter, distant cup clink"
}

Natural-language version:

Two friends sit across from each other in a quiet cafe. Medium two-shot with a static camera. Friend A looks up from the table, smiles, and says, “We made it.” Soft cafe chatter continues in the background with one distant cup clink.

Google recommends clear character description and explicitly attributed speech when prompting Veo for dialogue.

Tested output: The generated clip should keep Friend A as the only speaker, with the dialogue “We made it.” clearly attributed to Friend A. Check speaker consistency, lip synchronization, facial expression, and whether Friend B remains silent.

Revision to test: If speaker attribution drifts, remove the second character's movement and specify: “Only Friend A speaks. Friend B remains silent.”

3. Cinematic aerial

Use case: establishing shot

{
  "subject": "rocky coastline at sunrise",
  "action": "low clouds move over the cliffs while waves break below",
  "camera": "slow aerial rise and pullback",
  "lighting": "golden sunrise with cool blue shadows",
  "audio": "ocean surf and distant wind, no music"
}

Natural-language version:

A rocky coastline at sunrise. Low clouds move over the cliffs while waves break below. The camera slowly rises and pulls backward to reveal more of the coast. Golden sunlight falls across the rocks with cool blue shadows. Ocean surf and distant wind, no music.

Tested output: The generated clip should show a smooth, slow aerial rise and pullback over the coastline while maintaining stable cliffs, clouds, and waves. Check whether the camera movement remains consistent and whether the landscape avoids visible distortion.

Revision to test: If camera movement becomes too fast, change to: “Very slow aerial rise. Constant speed. No rotation.”

4. Food close-up

Use case: restaurant or food advertising

{
  "subject": "fresh ramen bowl on a dark wooden counter",
  "action": "steam rises while a spoon gently lifts one strand of noodles",
  "camera": "macro close-up with a very slow push-in",
  "lighting": "warm restaurant practical lights with soft highlights",
  "audio": "quiet room tone, subtle simmer, no dialogue"
}

Natural-language version:

A fresh ramen bowl sits on a dark wooden counter. Steam rises as a spoon gently lifts one strand of noodles. Macro close-up with a very slow camera push-in. Warm restaurant practical lights create soft highlights. Quiet room tone and a subtle simmer, no dialogue.

Tested output: The generated clip should maintain the ramen bowl's shape and appearance while steam rises naturally and the spoon lifts one strand of noodles. Check food consistency, steam behavior, spoon interaction, and camera stability.

Revision to test: If the food changes shape, remove the spoon interaction and keep only steam plus camera motion.

5. Vertical social ad

Use case: short product hook

{
  "subject": "white running shoe on wet pavement at dusk",
  "action": "the shoe lands once, creating a small realistic splash, then holds in frame",
  "camera": "quick low-angle tracking move that settles into a still close-up",
  "lighting": "warm storefront reflections against cool evening light",
  "audio": "single footstep, water splash, light city ambience"
}

Natural-language version:

A white running shoe lands once on wet pavement at dusk, creating a small realistic splash. A quick low-angle tracking move follows the landing and settles into a still close-up. Warm storefront reflections contrast with cool evening light. Audio includes one footstep, the splash, and light city ambience.

For vertical delivery, choose 9:16 in the interface rather than relying on a custom JSON field. Both Google and AKOOL currently document portrait support for Veo 3.1.

Tested output: The generated clip should show the running shoe landing once, producing a small splash, followed by a stable close-up. Check shoe geometry, splash realism, tracking-camera stability, and whether the scene remains correctly framed in 9:16.

Revision to test: If shoe geometry drifts, reduce the camera move and shorten the action to one impact plus hold.

How should you write dialogue and audio without common mistakes?

Keep dialogue short, identify the speaker clearly, and separate speech from ambience and sound effects.

Weak planning structure:

{
  "audio": "people talk, cafe sounds, music, someone says we made it"
}

Better:

{
  "dialogue": [
    {
      "speaker": "friend A",
      "line": "We made it."
    }
  ],
  "ambient_audio": "soft cafe chatter",
  "sound_effects": "one distant ceramic cup clink",
  "music": "none"
}

The benefit is not that Veo magically understands JSON better. The benefit is that you can see whether the brief contains conflicting audio instructions before converting it to prose.

Keep spoken lines short enough to fit the selected clip. Veo 3.1's current API supports clips up to 8 seconds, so a long paragraph of dialogue is not a realistic request for one generation.

Then test lip synchronization and speaker assignment in the actual output. Do not assume a correctly structured prompt guarantees exact timing.

JSON prompt vs plain text: which should you use?

Use JSON when structure and reuse help your workflow. Use plain text when you are prompting directly inside an interface that expects natural language. Neither format inherently produces better visual quality.

FactorJSON planning promptPlain-text prompt
ReadabilityStrong for structured teamsStrong for natural creative direction
ReuseEasy to template and storeEasy to copy and revise
Direct AKOOL UI supportNot publicly documented as a dedicated modeYes, normal prompting is documented
Syntax errorsPossibleNo JSON syntax to break
API useOnly when matching an actual API schemaUsually becomes the API's prompt value
Creative resultDepends on model interpretation after conversionDepends on model interpretation
Best usePlanning, automation, databases, prompt buildersDirect generation and iteration

A veo3 json prompt generator can therefore be useful as a front-end organizational tool. Its output should still be mapped to the fields or prompt syntax supported by the destination application.

JSON is structure, not a quality multiplier.

Why did my Veo JSON prompt fail?

Most failures come from invalid JSON, vague creative values, too many simultaneous actions, unsupported keys, or audio instructions that do not fit the clip.

ProblemExampleFix
Invalid JSON{"subject":"shoe",}Remove the trailing comma
Wrong quote marks{'subject':'shoe'}Use valid double quotes
Vague action"action":"cinematic"Describe physical movement
Vague camera"camera":"dynamic"Use “slow dolly in” or another real move
Too many eventsSix actions in one 8-second shotSplit into separate shots
Unsupported key"lens_magic":"epic"Remove it or convert the intention into prose
UI setting inside planning JSON"aspect_ratio":"9:16"Set 9:16 in the actual interface
Audio mismatch60-word dialogue in a short clipShorten the spoken line
Subject driftComplex product movementKeep product static and move the camera
JSON pasted literally into unsupported UIModel interprets braces as prompt textConvert field values into natural language

For example, this JSON is valid but creatively weak:

{
  "subject": "car",
  "action": "cool movement",
  "camera": "cinematic",
  "lighting": "nice"
}

Replace ambiguous values:

{
  "subject": "dark green vintage coupe on a coastal road",
  "action": "the car enters a left-hand curve at steady speed",
  "camera": "low front three-quarter tracking shot",
  "lighting": "late-afternoon sunlight with long warm reflections"
}

Then convert those values into a normal Veo prompt.

Frequently asked questions
How to use JSON prompt in Veo 3?
Does Veo 3.1 require JSON?
Can I paste JSON into AKOOL's Veo interface?
What fields should a Veo JSON prompt include?
Does a JSON prompt produce better video than a normal prompt?
AKOOL Content Team
Learn more
References

You may also like
No items found.
AKOOL Content Team