AI Video Generation
Text to video
Describe the shot in prompt and choose a text-to-video model. The clip length sets how long the generated video is.
{
"asset": {
"type": "video",
"prompt": "A slow aerial shot over a misty pine forest at sunrise, the camera drifting forward",
"model": "seedance-2.0-text-to-video",
"options": {
"aspectRatio": "16:9",
"resolution": "720p",
"generateAudio": false
}
},
"start": 0,
"length": 6
}
Image to video
An image-to-video model animates a still. Put the image URL in options.startSrc and describe the motion in prompt.
Both are required. Add options.endSrc to guide the final frame as well.
{
"asset": {
"type": "video",
"prompt": "Slowly zoom out and orbit left around the handbag",
"model": "seedance-2.0-image-to-video",
"options": {
"startSrc": "https://shotstack-assets.s3.amazonaws.com/images/handbag-flower-peaches.jpg",
"aspectRatio": "16:9"
}
},
"start": 0,
"length": 5
}
The image URL must be publicly reachable. It can also point to an image generated in the same render, using an alias.
Models
| Model | Length (seconds) | Resolution | Audio |
|---|---|---|---|
seedance-2.0-text-to-video, seedance-2.0-image-to-video | 4 to 15 | 480p, 720p (default), 1080p | Optional, off by default |
seedance-2.5-text-to-video, seedance-2.5-image-to-video | 4 to 30 | 480p, 720p (default), 1080p | Optional, off by default |
wan-3.0-prime-text-to-video, wan-3.0-prime-image-to-video | 2 to 30 | 480p, 720p, 1080p (default) | Optional, on by default |
gemini-omni-flash-1.1-text-to-video, gemini-omni-flash-1.1-image-to-video | 3 to 10 | 360p, 720p (default), 1080p, 4k | Always on |
Seedance 2.0 is the default. It costs less per second than Seedance 2.5, so choose 2.5 when one take needs to run longer than 15 seconds.
Options
| Option | Models | Values |
|---|---|---|
startSrc | Image-to-video models | URL of the first frame. Required. |
endSrc | Image-to-video models | URL of an image to guide the last frame. |
aspectRatio | Seedance | auto, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 |
| Wan | adaptive (default), 16:9, 4:3, 1:1, 3:4, 9:16 | |
| Gemini Omni Flash | 16:9 (default), 9:16 | |
resolution | All | See the table above. |
generateAudio | Seedance, Wan | true or false. Generates sound with the picture. |
enhancePrompt | Wan | true (default) lets the model expand your prompt. |
seed | Wan | An integer. The same seed and prompt reproduce a result. |
GET /models/{id} returns the full option schema for a model.
Match aspectRatio to your output size. A 9:16 video in a 16:9 edit is cropped to fit, and the cropped pixels are paid
for.
Length
The clip length sets the duration of the generated video, rounded up to a whole second. A clip shorter than the
model's minimum gets a video of the minimum length, trimmed to the clip. A clip longer than the maximum gets a video of
the maximum length, which ends before the clip does, so keep video clips within the model's range.
With "length": "auto", the model makes its default length (15 seconds for Seedance, 5 for Wan, 8 for Gemini Omni Flash)
and the clip takes the length of the generated file. Give video clips a numeric length when you know it: the price
follows the length, and without one the quote is an estimate.
Video generation takes minutes. A render waits for it with a status of generating.
Writing video prompts
This guidance is for Seedance, the most widely used of these models.
- Spend the words on motion. Describe what moves and what the movement causes ("leaves scatter on each impact") rather than how things look at rest.
- Fit the action to the length. Packing several events into a short clip makes it rushed and incoherent. A prompt written for 15 seconds won't hold up at 5.
- Ask for cuts with "cut to". A single generation stays consistent across about three cuts, which holds together better than the same shots generated separately.
- Silence is an option, not a phrase. Set
generateAudiotofalse. Writing "no music" in the prompt doesn't work. - Dialogue in double quotes is lip-synced when audio is on. Say what tone of voice you want.
- Don't repeat settings in the prompt. Writing the aspect ratio or resolution in the prompt does nothing; set the option.
When animating an image, the image decides the composition, lighting and style, so the prompt should describe only the motion. A badly composed or compressed start image can't be fixed in the prompt, and its flaws get worse once it moves. Chaining clips by feeding one clip's last frame into the next drifts quickly. Keep chains to two or three clips.