Skip to main content

AI Speech Generation

Text to speech​

Put the script in prompt, choose a speech model and pick a voice. Give the clip "length": "auto" so it takes the length of the generated speech.

{
"asset": {
"type": "audio",
"prompt": "Good evening. Storms are rolling in across Sydney tonight, bringing flash flood warnings and a slow commute home.",
"model": "elevenlabs-multilingual-v2",
"options": {
"voice": "George"
}
},
"start": 0,
"length": "auto"
}

Speech length can't be known before it is generated, so auto is better than guessing a length. See smart clips for placing other clips around it.

Models​

ModelUse it forVoices
elevenlabs-multilingual-v2The default. Natural, consistent narration in 29 languages.ElevenLabs voices, Rachel by default
elevenlabs-turbo-v2.5Faster generation in 32 languages.ElevenLabs voices, Rachel by default
minimax-speech-2.8-hdNarration and spoken stories.MiniMax voices, Wise_Woman by default
polly-neuralPlain narration and news-style reads.Amazon Polly neural voices. voice is required.

Options​

OptionModelsValues
voiceElevenLabsRachel, Aria, Roger, Sarah, Laura, Charlie, George, Callum, River, Liam, Charlotte, Alice, Matilda, Will, Jessica, Eric, Chris, Brian, Daniel, Lily, Bill
MiniMaxWise_Woman, Friendly_Person, Inspirational_girl, Deep_Voice_Man, Calm_Woman, Casual_Guy, Lively_Girl, Patient_Man, Young_Knight, Determined_Man, Lovely_Girl, Decent_Boy, Imposing_Manner, Elegant_Man, Abbess, Sweet_Girl_2, Exuberant_Girl
PollyA Polly neural voice name, such as Joanna or Olivia
languageElevenLabsA two-letter language code, such as en or de
MiniMaxA language name, such as English, Japanese or Chinese,Yue, or auto
PollyA language code, such as en-US or es-MX
stabilityElevenLabs0 to 1. Lower is more expressive, higher is more even.
similarityBoostElevenLabs0 to 1. How closely the output keeps to the original voice.
styleElevenLabs0 to 1. How strongly the voice's style is exaggerated.
newscasterPollytrue for a newsreader delivery. Works with Matthew, Joanna, Lupe and Amy only.

GET /models/{id} returns the full option schema for a model. Write the prompt in the language you want spoken.

Writing scripts​

The script does most of the work. Most poor results are a writing problem, not a settings problem.

  • Spell out what must be read a certain way. Numbers, currency, dates and abbreviations are the most common mistakes.
  • Write only what should be spoken. Audio tags like [laughs] have no effect on these models, and stage directions written into the script are read aloud.
  • Choose a voice by listening. Voice names say little about how a voice sounds. Generate a short line with a few voices and compare.
  • Newscaster on other voices fails. Setting newscaster on a Polly voice other than the four listed fails the generation.

Captions​

Generated speech can be captioned in the same render. Give the speech clip an alias and point a caption clip at it. See rich captions and aliases.