What You Will Build
Synchronized Sound
Ambient audio and effects that match the action on screen
Spoken Lines
Dialogue or narration described in the prompt, voiced by the model
One Request
Audio is generated together with the video, in the same job
Predictable Cost
A fixed with-audio rate per second, shown before you generate
Before You Begin
Make sure you have:- A Pictory API key
- Node.js or Python installed on your machine
- The required packages installed
Which Models Support Audio
omni-flash always produces audio. You do not need to send audio: true for it, and audio: false is rejected with a 400 validation error. Every other model generates a silent video unless you set audio: true.Step-by-Step Guide
Step 1: Set Up Your Request
Addaudio: true to an otherwise ordinary video request. Describe the sounds you expect in the prompt: the model uses the description to decide what to voice.
audio can be combined with every other option on the endpoint: reference images, a first frame, a first and end frame transition, or a video to extend. When extending on veo3.1 or veo3.1_fast, audio applies to the new segment.Step 2: Submit the Video Generation Request
Send the request to the AI Studio video generation endpoint.Step 3: Poll for the Result
Check the job status at regular intervals until the video is ready. The completed response is the same as for a silent video; the audio track is inside the delivered MP4.Understanding the Parameters
How Audio Is Billed
Audio is not a surcharge on top of the silent rate; it is a separate per-second rate for the model. A8s video on pixverse6 costs 21.6 credits without audio and 29.6 credits with audio. Credits are reserved when the job is created and released if the generation fails. The aiCreditsUsed field on the completed job shows the final charge.
Tips for Video with Audio
- Describe the soundscape. Mention the sounds you want (“rain on a tin roof”, “a crowd cheering”) as well as the visuals. The model voices what the prompt describes.
- Quote the dialogue. For speech, put the exact words in the prompt, for example: a barista says “Your coffee is ready.”
- Keep clips self-contained. Audio is generated per job, so when you build multi-segment videos, describe a natural sound at the end of one segment and the start of the next to avoid abrupt cuts.
- Check the listing. Videos returned by Get Generated Videos carry an
audioflag so you can tell silent and sound clips apart later.
Next Steps
- Generate Video from Text Prompt for the basic video workflow
- Generate Video from Reference Images to keep subjects consistent
- Extend Video with AI to continue an existing video
- Generate Video API Reference for the complete parameter documentation
