Generating Assets with AI
Video, image and audio assets can be generated from a text prompt. Set prompt on the asset and choose a model, and
Shotstack generates the file and places it in the clip. Generation runs on your Shotstack API key and is paid for in
credits, so there is no separate account to open with each AI provider and nothing extra to integrate.
{
"asset": {
"type": "video",
"prompt": "A slow aerial shot over a misty pine forest at sunrise",
"model": "seedance-2.0-text-to-video",
"options": {
"resolution": "720p",
"aspectRatio": "16:9",
"generateAudio": false
}
},
"start": 0,
"length": 5
}
Each kind of media has its own guide:
- Video: generate from a prompt, or animate an image.
- Images: generate from a prompt, or edit existing images.
- Speech: turn a script into a voiceover.
- Music: generate a track from a prompt or a section-by-section plan.
Pricing and quotes explains what a generation costs and how to check the price before you generate.
Two ways to generate
Inside a render. Put prompt assets in an edit and POST it to the render endpoint as usual. The render status is
generating while the assets are made, then the render carries on with the generated files in place. Use this when the
media only exists to be part of the edit.
On its own. POST a single asset to the generate endpoint and you get back a URL to the generated file. Use this to
check the output before you render, to reuse one asset across many edits, or to use it outside Shotstack.
Both use the same models and prices, and both draw on the same stored generations. An image you generated with the
generate endpoint is reused at no charge by a render that asks for the same image, and the other way round.
The generate endpoint
The request body is an asset, written exactly as it would be inside a clip:
curl -X POST \
-H "Content-Type: application/json" \
-H "x-api-key: $SHOTSTACK_API_KEY" \
-d '{"asset": {"type": "image", "prompt": "A lighthouse on a rocky coast at sunset, cinematic lighting", "model": "nano-banana-2"}}' \
https://api.shotstack.io/edit/stage/generate
If the same asset has been generated before, the response is 200 with the file straight away:
{
"id": "8a1f2c3d-4e5b-5a6c-9d7e-1f2a3b4c5d6e",
"status": "done",
"url": "https://cdn.shotstack.io/au/stage/generations/f0a9f2b1c3/8a1f2c3d-4e5b-5a6c-9d7e-1f2a3b4c5d6e.png"
}
Otherwise the response is 202 with a status of queued. Poll the job until it finishes:
curl -X GET \
-H "x-api-key: $SHOTSTACK_API_KEY" \
https://api.shotstack.io/edit/stage/generate/8a1f2c3d-4e5b-5a6c-9d7e-1f2a3b4c5d6e
While it runs, the response is 202 with a status of processing and a Retry-After header giving the seconds to wait.
When it ends, the response is 200 with either done and a url, or failed and an error. Most images and speech
take seconds. Video, music and the GPT Image models can take minutes.
For video and music, add length beside asset to set the duration in seconds:
{
"asset": {
"type": "video",
"prompt": "Waves rolling onto a black sand beach, low angle",
"model": "seedance-2.0-text-to-video"
},
"length": 8
}
The API reference lists every response and error.
Choosing a model
GET /models lists every model you can generate with. Build pickers from it rather than hard-coding model names, so new
models appear without a code change.
curl -X GET \
-H "x-api-key: $SHOTSTACK_API_KEY" \
"https://api.shotstack.io/edit/stage/models?expand=options"
Each entry has the model id to put on the asset, its type, a name and description, and, with
expand=options, a JSON Schema for its options. available says whether your plan can use the model. When it is
false, unavailableReason says why. GET /models/{id} returns a single model with its options.
When model is left out, the asset uses the default for its type:
| Asset type | Default model |
|---|---|
image | nano-banana-2 |
video | seedance-2.0-text-to-video |
audio | elevenlabs-multilingual-v2 (speech) |
Set the model explicitly. An audio prompt with no model is read aloud as speech, so music needs
"model": "elevenlabs-music". Animating an image needs an image-to-video model, because the default video model works
from the prompt alone.
options holds settings for the chosen model only. An option the model doesn't accept, or a value outside its schema, is
rejected with an error that names the option.
When an asset is generated
What happens to an asset depends on whether it has a src, a prompt or both:
| The asset has | What happens |
|---|---|
src only | The file at src is used. Nothing is generated or charged. |
prompt only | The asset is generated, or reused if it was generated before. |
src and prompt | The prompt wins. src is treated as a preview placeholder and replaced by the generated file. |
The third case lets an editor show a preview while the edit still describes how to make the asset. It never costs twice:
an unchanged prompt resolves to the file it produced last time. To fix an asset so it can't change again, delete the
prompt and keep the src.
Reusing earlier generations
Every generated file is stored against your account. When an asset asks for exactly what was generated before, the stored file is used and nothing is charged, so you never pay twice for the same asset.
An asset counts as the same when all of these match:
- The asset
typeand themodel. Leavingmodelout is the same as naming the default. - The
prompt, ignoring spaces at either end. Any other change, including case and punctuation, is a new prompt. - The
optionsyou set. - For video and music, the length: the clip
lengthin a render, orlengthon thegenerateendpoint. - For video, the start and end image URLs, ignoring the query string, so a re-signed URL to the same image still matches.
Reference images for image editing (
imageUrls) must match exactly, query string included.
There is no switch to force a fresh generation of an identical asset. To get a different result, change the prompt, or
set a different seed on a model that accepts one.
The sandbox and production keep separate stores. An asset generated in the sandbox is generated, and charged, again the
first time production asks for it. To carry one across, generate it once with the generate endpoint and use the
returned URL as the src of an asset with no prompt.
Rendering with generated assets
- Failures. If an asset fails to generate, the render fails and the error names the reason. Assets that did generate are kept, so rendering again reuses them.
- Merge fields. Placeholders in
promptandoptionsare replaced before generation, so a template can generate a different asset for each render. Thegenerateendpoint needs final values. - Clip length. For video and music, the clip
lengthsets how long the generated media is. With"auto", the model makes its default length and the clip takes the length of the file it produced. Speech has no length setting: give it"length": "auto"and let the clip fit the speech. See smart clips. - Captions. Generated speech can be captioned automatically by pointing a caption clip at it with an alias.
Animating an image generated in the same render
A video's startSrc or endSrc can point to another generated clip with an alias.
The image is generated first and its file is passed to the video model. The image clip must have a prompt and no src.
{
"timeline": {
"tracks": [
{
"clips": [
{
"alias": "product",
"asset": {
"type": "image",
"prompt": "A ceramic coffee cup on a wooden table, soft morning light",
"model": "nano-banana-2",
"options": { "aspectRatio": "16:9", "resolution": "2K" }
},
"start": 0,
"length": 3
},
{
"asset": {
"type": "video",
"prompt": "Steam rises from the cup as the camera slowly pushes in",
"model": "seedance-2.0-image-to-video",
"options": { "startSrc": "alias://product", "aspectRatio": "16:9" }
},
"start": 3,
"length": 5
}
]
}
]
},
"output": {
"format": "mp4",
"size": { "width": 1280, "height": 720 }
}
}
Storing generated media
Generated files are kept until you delete them. The generate endpoint returns a permanent URL on cdn.shotstack.io
that you can use in any edit or outside Shotstack. Files generated during renders are stored too.
Generated media is listed under My Generations in the dashboard, where you can preview a file, copy its URL or delete it. A deleted file is gone for good: the next request for the same asset generates and charges again.
Using your own provider
Generating through Shotstack is a convenience, not a requirement. If you already work with an AI provider, or need a
model that isn't listed, generate the file there and give its URL to an asset as src, with no prompt. It renders like
any other media, and Shotstack charges nothing to generate it.
Legacy asset types
The text-to-image, text-to-speech and image-to-video asset types are deprecated. Edits that use them keep working:
Shotstack converts them to the asset types above. The video models behind image-to-video are deprecated too, so move
those edits to a current video model. For new work, use the
current form:
| Legacy asset | Current equivalent |
|---|---|
text-to-image with prompt | image with prompt and an image model |
text-to-speech with text and voice | audio with the text as prompt, a speech model, and voice in options |
image-to-video with src and prompt | video with prompt, an image-to-video model, and the image as options.startSrc |