Skip to main content
POST
Generate Video

Overview

The Generate Video API creates an AI-generated video based on a text prompt. You can choose from multiple AI video models, specify an aspect ratio and duration, generate synchronized audio, and optionally provide a first frame image (with an optional end frame for a transition), a video to extend, or reference images to guide the output. The API returns a job ID that you can use to poll for the result once the video generation is complete.
A valid API key is required to use this endpoint. Obtain your API key from the API Access page in your Pictory dashboard.

API Endpoint


Request Headers

string
required
API key for authentication (starts with pictai_)
string
required
Must be application/json

Request Body

string
required
A text description of the video you want to generate. The prompt must be between 5 and 5,000 characters.Example: "A drone flying over a lush green forest with sunlight filtering through the canopy"
string
The AI model to use for video generation. Each model offers different quality levels, durations, aspect ratios, and features. Defaults to pixverse5.6 if not specified.Supported models: pixverse6, pixverse5.6, veo3.1, veo3.1_fast, omni-flashModel Capabilities and Pricing:Feature Support by Model:
pixverse5.6 is the budget option: the same durations and aspect ratios as before at the lowest rate. pixverse6 adds 15s, wide and portrait ratios such as 21:9 and 2:3, and video extension. Requests that still send the retired pixverse5.5 are served and billed as pixverse5.6; validation still runs against the pixverse5.5 entry, so send pixverse5.6 to use audio.
string
The aspect ratio for the generated video. The available values depend on the selected model. Defaults to 16:9. Refer to the model table above for supported aspect ratios per model. An extension (extendVideoUrl) or a transition (firstFrameImageUrl with endFrameImageUrl) follows the source video or the frames, so the field has no effect there.
string
The duration of the generated video. The available values depend on the selected model. Defaults to the shortest duration of the chosen model (5s on Pixverse, 4s on Veo, 10s on omni-flash). Refer to the model table above for supported durations per model.Two situations narrow the list:
  • Reference images on Veo: veo3.1 and veo3.1_fast generate 8s videos when referenceImageUrls is provided. Any other value is rejected, and the default becomes 8s.
  • Extending on Veo: the continuation is a fixed 7s, so 7s is the only accepted value and the default.
boolean
Generate synchronized audio (ambient sound, effects, and speech implied by the prompt) with the video. Defaults to false. Audio is billed at the model’s “With Audio” rate shown in the table above, per second of video.omni-flash always produces audio and cannot be muted: the field defaults to true for that model, and sending audio: false is rejected with a 400 validation error.
Setting audio: true on a model that does not support audio returns a 400 validation error.
string
An optional URL of an image to use as the first frame of the video. This guides the visual starting point for the generation. The URL must be a valid, publicly accessible URI.
This parameter cannot be used together with extendVideoUrl or referenceImageUrls.
string
An optional URL of an image the video should end on. Requires firstFrameImageUrl: the model generates a transition that starts on the first frame and ends on this one. Supported by pixverse6 and pixverse5.6. The URL must be a valid, publicly accessible URI.This is the end frame you supply as input. The lastFrameImageUrl returned on completed jobs is the last frame extracted from the generated video, which you can feed into a follow-up request.
Sending endFrameImageUrl without firstFrameImageUrl, or on a model without first and end frame support, returns a 400 validation error.
string
An optional URL of an existing video to extend. The AI model will generate additional content that continues from the end of the provided video. The URL must be a valid, publicly accessible http or https URI. Supported by pixverse6, veo3.1, and veo3.1_fast; the default model for an extension request is pixverse6.On veo3.1 and veo3.1_fast the continuation is a fixed 7s, and the delivered file contains the source video followed by the new segment. On pixverse6 the duration you request is the length of the new segment.
This parameter cannot be used together with firstFrameImageUrl, endFrameImageUrl, or referenceImageUrls.
array
An optional array of image URLs to use as visual references for the generated video. The AI model uses these images to keep subjects, objects, or styles consistent in the output. Each URL must be a valid, publicly accessible URI.The maximum number of references depends on the model: up to 4 on pixverse6 and pixverse5.6, up to 3 on veo3.1 and veo3.1_fast, and 1 on omni-flash. Sending more than the model allows returns a 400 validation error.On pixverse6 and pixverse5.6 you can name each reference in the prompt as @subject1, @subject2, and so on, in array order. References you do not mention are appended to the prompt automatically.
This parameter cannot be used together with firstFrameImageUrl or extendVideoUrl. On veo3.1 and veo3.1_fast, reference images require an 8s duration.
string
Steer the generation toward a training-content use case. The preset is a fixed directive applied during prompt enhancement, so sending it always turns enhancement on, even when enhancePrompt is false. Supported values: sales-training, product-training, soft-skills-training.
boolean
Whether Pictory rewrites your prompt with an AI enhancer before generation, adding cinematic and compositional detail. Off by default: your prompt goes to the model exactly as written. Set to true to enhance it. Sending a trainingPreset always turns enhancement on, even when enhancePrompt is false. When enhancement runs, the completed job returns the rewritten text as enhancedPrompt.
string
Output resolution: "1080p" (default) or "720p". Without it, you get the highest resolution the model supports. The model delivers the highest it supports up to this value: omni-flash is always 720p, pixverse6 at 15s is 720p, and Veo is 1080p at 8s only. Credits do not change with resolution.
string
An optional webhook URL that will receive a notification when the video generation job completes. The URL must be a valid URI.

Response

Success Response (200)

boolean
true when the job has been created successfully
object
string
The unique identifier (UUID) of the video generation job. Use this ID to poll for the result using the Get Video Generation Job endpoint.

Response Examples


Status Codes


Output Resolution

Every video is delivered at the highest resolution its model supports; there is no quality parameter to set. For 1080p that is 1920×1080 for 16:9, 1080×1920 for 9:16, and the matching 1080-pixel size for the other ratios. The Get Model Capabilities endpoint reports maxResolution per model. The width and height on the completed job always report the delivered size. Credits do not depend on resolution.

Parameter Exclusivity

The firstFrameImageUrl, extendVideoUrl, and referenceImageUrls parameters are mutually exclusive. You may use only one of these in a single request. Providing more than one will result in a 400 validation error. endFrameImageUrl is the one addition: it is only valid alongside firstFrameImageUrl. audio, trainingPreset, and enhancePrompt can be combined with any of them.

Code Examples

Replace YOUR_API_KEY with your actual API key from the API Access page.

Next Steps

After receiving the jobId, poll for the video generation result using the Get Video Generation Job endpoint. Use a polling interval of 10–30 seconds to check the job status.