Skip to main content

AI Video Generation

Text to video​

Describe the shot in prompt and choose a text-to-video model. The clip length sets how long the generated video is.

{
"asset": {
"type": "video",
"prompt": "A slow aerial shot over a misty pine forest at sunrise, the camera drifting forward",
"model": "seedance-2.0-text-to-video",
"options": {
"aspectRatio": "16:9",
"resolution": "720p",
"generateAudio": false
}
},
"start": 0,
"length": 6
}

Image to video​

An image-to-video model animates a still. Put the image URL in options.startSrc and describe the motion in prompt. Both are required. Add options.endSrc to guide the final frame as well.

{
"asset": {
"type": "video",
"prompt": "Slowly zoom out and orbit left around the handbag",
"model": "seedance-2.0-image-to-video",
"options": {
"startSrc": "https://shotstack-assets.s3.amazonaws.com/images/handbag-flower-peaches.jpg",
"aspectRatio": "16:9"
}
},
"start": 0,
"length": 5
}

The image URL must be publicly reachable. It can also point to an image generated in the same render, using an alias.

Models​

ModelLength (seconds)ResolutionAudio
seedance-2.0-text-to-video, seedance-2.0-image-to-video4 to 15480p, 720p (default), 1080pOptional, off by default
seedance-2.5-text-to-video, seedance-2.5-image-to-video4 to 30480p, 720p (default), 1080pOptional, off by default
wan-3.0-prime-text-to-video, wan-3.0-prime-image-to-video2 to 30480p, 720p, 1080p (default)Optional, on by default
gemini-omni-flash-1.1-text-to-video, gemini-omni-flash-1.1-image-to-video3 to 10360p, 720p (default), 1080p, 4kAlways on

Seedance 2.0 is the default. It costs less per second than Seedance 2.5, so choose 2.5 when one take needs to run longer than 15 seconds.

Options​

OptionModelsValues
startSrcImage-to-video modelsURL of the first frame. Required.
endSrcImage-to-video modelsURL of an image to guide the last frame.
aspectRatioSeedanceauto, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16
Wanadaptive (default), 16:9, 4:3, 1:1, 3:4, 9:16
Gemini Omni Flash16:9 (default), 9:16
resolutionAllSee the table above.
generateAudioSeedance, Wantrue or false. Generates sound with the picture.
enhancePromptWantrue (default) lets the model expand your prompt.
seedWanAn integer. The same seed and prompt reproduce a result.

GET /models/{id} returns the full option schema for a model.

Match aspectRatio to your output size. A 9:16 video in a 16:9 edit is cropped to fit, and the cropped pixels are paid for.

Length​

The clip length sets the duration of the generated video, rounded up to a whole second. A clip shorter than the model's minimum gets a video of the minimum length, trimmed to the clip. A clip longer than the maximum gets a video of the maximum length, which ends before the clip does, so keep video clips within the model's range.

With "length": "auto", the model makes its default length (15 seconds for Seedance, 5 for Wan, 8 for Gemini Omni Flash) and the clip takes the length of the generated file. Give video clips a numeric length when you know it: the price follows the length, and without one the quote is an estimate.

Video generation takes minutes. A render waits for it with a status of generating.

Writing video prompts​

This guidance is for Seedance, the most widely used of these models.

  • Spend the words on motion. Describe what moves and what the movement causes ("leaves scatter on each impact") rather than how things look at rest.
  • Fit the action to the length. Packing several events into a short clip makes it rushed and incoherent. A prompt written for 15 seconds won't hold up at 5.
  • Ask for cuts with "cut to". A single generation stays consistent across about three cuts, which holds together better than the same shots generated separately.
  • Silence is an option, not a phrase. Set generateAudio to false. Writing "no music" in the prompt doesn't work.
  • Dialogue in double quotes is lip-synced when audio is on. Say what tone of voice you want.
  • Don't repeat settings in the prompt. Writing the aspect ratio or resolution in the prompt does nothing; set the option.

When animating an image, the image decides the composition, lighting and style, so the prompt should describe only the motion. A badly composed or compressed start image can't be fixed in the prompt, and its flaws get worse once it moves. Chaining clips by feeding one clip's last frame into the next drifts quickly. Keep chains to two or three clips.