Skip to main content

Generating Assets with AI

Video, image and audio assets can be generated from a text prompt. Set prompt on the asset and choose a model, and Shotstack generates the file and places it in the clip. Generation runs on your Shotstack API key and is paid for in credits, so there is no separate account to open with each AI provider and nothing extra to integrate.

{
"asset": {
"type": "video",
"prompt": "A slow aerial shot over a misty pine forest at sunrise",
"model": "seedance-2.0-text-to-video",
"options": {
"resolution": "720p",
"aspectRatio": "16:9",
"generateAudio": false
}
},
"start": 0,
"length": 5
}

Each kind of media has its own guide:

  • Video: generate from a prompt, or animate an image.
  • Images: generate from a prompt, or edit existing images.
  • Speech: turn a script into a voiceover.
  • Music: generate a track from a prompt or a section-by-section plan.

Pricing and quotes explains what a generation costs and how to check the price before you generate.

Two ways to generate​

Inside a render. Put prompt assets in an edit and POST it to the render endpoint as usual. The render status is generating while the assets are made, then the render carries on with the generated files in place. Use this when the media only exists to be part of the edit.

On its own. POST a single asset to the generate endpoint and you get back a URL to the generated file. Use this to check the output before you render, to reuse one asset across many edits, or to use it outside Shotstack.

Both use the same models and prices, and both draw on the same stored generations. An image you generated with the generate endpoint is reused at no charge by a render that asks for the same image, and the other way round.

The generate endpoint​

The request body is an asset, written exactly as it would be inside a clip:

curl -X POST \
-H "Content-Type: application/json" \
-H "x-api-key: $SHOTSTACK_API_KEY" \
-d '{"asset": {"type": "image", "prompt": "A lighthouse on a rocky coast at sunset, cinematic lighting", "model": "nano-banana-2"}}' \
https://api.shotstack.io/edit/stage/generate

If the same asset has been generated before, the response is 200 with the file straight away:

{
"id": "8a1f2c3d-4e5b-5a6c-9d7e-1f2a3b4c5d6e",
"status": "done",
"url": "https://cdn.shotstack.io/au/stage/generations/f0a9f2b1c3/8a1f2c3d-4e5b-5a6c-9d7e-1f2a3b4c5d6e.png"
}

Otherwise the response is 202 with a status of queued. Poll the job until it finishes:

curl -X GET \
-H "x-api-key: $SHOTSTACK_API_KEY" \
https://api.shotstack.io/edit/stage/generate/8a1f2c3d-4e5b-5a6c-9d7e-1f2a3b4c5d6e

While it runs, the response is 202 with a status of processing and a Retry-After header giving the seconds to wait. When it ends, the response is 200 with either done and a url, or failed and an error. Most images and speech take seconds. Video, music and the GPT Image models can take minutes.

For video and music, add length beside asset to set the duration in seconds:

{
"asset": {
"type": "video",
"prompt": "Waves rolling onto a black sand beach, low angle",
"model": "seedance-2.0-text-to-video"
},
"length": 8
}

The API reference lists every response and error.

Choosing a model​

GET /models lists every model you can generate with. Build pickers from it rather than hard-coding model names, so new models appear without a code change.

curl -X GET \
-H "x-api-key: $SHOTSTACK_API_KEY" \
"https://api.shotstack.io/edit/stage/models?expand=options"

Each entry has the model id to put on the asset, its type, a name and description, and, with expand=options, a JSON Schema for its options. available says whether your plan can use the model. When it is false, unavailableReason says why. GET /models/{id} returns a single model with its options.

When model is left out, the asset uses the default for its type:

Asset typeDefault model
imagenano-banana-2
videoseedance-2.0-text-to-video
audioelevenlabs-multilingual-v2 (speech)

Set the model explicitly. An audio prompt with no model is read aloud as speech, so music needs "model": "elevenlabs-music". Animating an image needs an image-to-video model, because the default video model works from the prompt alone.

options holds settings for the chosen model only. An option the model doesn't accept, or a value outside its schema, is rejected with an error that names the option.

When an asset is generated​

What happens to an asset depends on whether it has a src, a prompt or both:

The asset hasWhat happens
src onlyThe file at src is used. Nothing is generated or charged.
prompt onlyThe asset is generated, or reused if it was generated before.
src and promptThe prompt wins. src is treated as a preview placeholder and replaced by the generated file.

The third case lets an editor show a preview while the edit still describes how to make the asset. It never costs twice: an unchanged prompt resolves to the file it produced last time. To fix an asset so it can't change again, delete the prompt and keep the src.

Reusing earlier generations​

Every generated file is stored against your account. When an asset asks for exactly what was generated before, the stored file is used and nothing is charged, so you never pay twice for the same asset.

An asset counts as the same when all of these match:

  • The asset type and the model. Leaving model out is the same as naming the default.
  • The prompt, ignoring spaces at either end. Any other change, including case and punctuation, is a new prompt.
  • The options you set.
  • For video and music, the length: the clip length in a render, or length on the generate endpoint.
  • For video, the start and end image URLs, ignoring the query string, so a re-signed URL to the same image still matches. Reference images for image editing (imageUrls) must match exactly, query string included.

There is no switch to force a fresh generation of an identical asset. To get a different result, change the prompt, or set a different seed on a model that accepts one.

The sandbox and production keep separate stores. An asset generated in the sandbox is generated, and charged, again the first time production asks for it. To carry one across, generate it once with the generate endpoint and use the returned URL as the src of an asset with no prompt.

Rendering with generated assets​

  • Failures. If an asset fails to generate, the render fails and the error names the reason. Assets that did generate are kept, so rendering again reuses them.
  • Merge fields. Placeholders in prompt and options are replaced before generation, so a template can generate a different asset for each render. The generate endpoint needs final values.
  • Clip length. For video and music, the clip length sets how long the generated media is. With "auto", the model makes its default length and the clip takes the length of the file it produced. Speech has no length setting: give it "length": "auto" and let the clip fit the speech. See smart clips.
  • Captions. Generated speech can be captioned automatically by pointing a caption clip at it with an alias.

Animating an image generated in the same render​

A video's startSrc or endSrc can point to another generated clip with an alias. The image is generated first and its file is passed to the video model. The image clip must have a prompt and no src.

{
"timeline": {
"tracks": [
{
"clips": [
{
"alias": "product",
"asset": {
"type": "image",
"prompt": "A ceramic coffee cup on a wooden table, soft morning light",
"model": "nano-banana-2",
"options": { "aspectRatio": "16:9", "resolution": "2K" }
},
"start": 0,
"length": 3
},
{
"asset": {
"type": "video",
"prompt": "Steam rises from the cup as the camera slowly pushes in",
"model": "seedance-2.0-image-to-video",
"options": { "startSrc": "alias://product", "aspectRatio": "16:9" }
},
"start": 3,
"length": 5
}
]
}
]
},
"output": {
"format": "mp4",
"size": { "width": 1280, "height": 720 }
}
}

Storing generated media​

Generated files are kept until you delete them. The generate endpoint returns a permanent URL on cdn.shotstack.io that you can use in any edit or outside Shotstack. Files generated during renders are stored too.

Generated media is listed under My Generations in the dashboard, where you can preview a file, copy its URL or delete it. A deleted file is gone for good: the next request for the same asset generates and charges again.

Using your own provider​

Generating through Shotstack is a convenience, not a requirement. If you already work with an AI provider, or need a model that isn't listed, generate the file there and give its URL to an asset as src, with no prompt. It renders like any other media, and Shotstack charges nothing to generate it.

Legacy asset types​

The text-to-image, text-to-speech and image-to-video asset types are deprecated. Edits that use them keep working: Shotstack converts them to the asset types above. The video models behind image-to-video are deprecated too, so move those edits to a current video model. For new work, use the current form:

Legacy assetCurrent equivalent
text-to-image with promptimage with prompt and an image model
text-to-speech with text and voiceaudio with the text as prompt, a speech model, and voice in options
image-to-video with src and promptvideo with prompt, an image-to-video model, and the image as options.startSrc