AI Speech Generation
Text to speech
Put the script in prompt, choose a speech model and pick a voice. Give the clip "length": "auto" so it takes the
length of the generated speech.
{
"asset": {
"type": "audio",
"prompt": "Good evening. Storms are rolling in across Sydney tonight, bringing flash flood warnings and a slow commute home.",
"model": "elevenlabs-multilingual-v2",
"options": {
"voice": "George"
}
},
"start": 0,
"length": "auto"
}
Speech length can't be known before it is generated, so auto is better than guessing a length. See
smart clips for placing other clips around it.
Models
| Model | Use it for | Voices |
|---|---|---|
elevenlabs-multilingual-v2 | The default. Natural, consistent narration in 29 languages. | ElevenLabs voices, Rachel by default |
elevenlabs-turbo-v2.5 | Faster generation in 32 languages. | ElevenLabs voices, Rachel by default |
minimax-speech-2.8-hd | Narration and spoken stories. | MiniMax voices, Wise_Woman by default |
polly-neural | Plain narration and news-style reads. | Amazon Polly neural voices. voice is required. |
Options
| Option | Models | Values |
|---|---|---|
voice | ElevenLabs | Rachel, Aria, Roger, Sarah, Laura, Charlie, George, Callum, River, Liam, Charlotte, Alice, Matilda, Will, Jessica, Eric, Chris, Brian, Daniel, Lily, Bill |
| MiniMax | Wise_Woman, Friendly_Person, Inspirational_girl, Deep_Voice_Man, Calm_Woman, Casual_Guy, Lively_Girl, Patient_Man, Young_Knight, Determined_Man, Lovely_Girl, Decent_Boy, Imposing_Manner, Elegant_Man, Abbess, Sweet_Girl_2, Exuberant_Girl | |
| Polly | A Polly neural voice name, such as Joanna or Olivia | |
language | ElevenLabs | A two-letter language code, such as en or de |
| MiniMax | A language name, such as English, Japanese or Chinese,Yue, or auto | |
| Polly | A language code, such as en-US or es-MX | |
stability | ElevenLabs | 0 to 1. Lower is more expressive, higher is more even. |
similarityBoost | ElevenLabs | 0 to 1. How closely the output keeps to the original voice. |
style | ElevenLabs | 0 to 1. How strongly the voice's style is exaggerated. |
newscaster | Polly | true for a newsreader delivery. Works with Matthew, Joanna, Lupe and Amy only. |
GET /models/{id} returns the full option schema for a model. Write the prompt in the language you want spoken.
Writing scripts
The script does most of the work. Most poor results are a writing problem, not a settings problem.
- Spell out what must be read a certain way. Numbers, currency, dates and abbreviations are the most common mistakes.
- Write only what should be spoken. Audio tags like
[laughs]have no effect on these models, and stage directions written into the script are read aloud. - Choose a voice by listening. Voice names say little about how a voice sounds. Generate a short line with a few voices and compare.
- Newscaster on other voices fails. Setting
newscasteron a Polly voice other than the four listed fails the generation.
Captions
Generated speech can be captioned in the same render. Give the speech clip an alias and point a caption clip at it.
See rich captions and
aliases.