Short video hooks
The first three seconds of a vertical video, the use our note names. Write the hook as its own line.

Sienna carries the published list's only target quality of A among the American English voices, with 10 to 100 hours of training audio and an overall A-. Our note: brighter than Maya, good for hooks and short promos.
Hear it
On the record
Target quality measures how clean the reference audio was and how well its text matched; Sienna's is A. Her training band is 10 to 100 hours, the largest on the list; together they give the A-.
Our note places her next to Maya as the brighter of the two and points her at hooks and promos; keep hooks short.
Length the engine likes. The published guidance puts the sweet spot at roughly 100 to 200 tokens, a few sentences; very short lines are the weak spot and very long ones can rush.
Where we would use it
The first three seconds of a vertical video, the use our note names. Write the hook as its own line.
Launches, sales, and event reminders where the copy is short. Add a pause tag before the offer.
An even split with Brooke is the one Sienna blend we kept. Start there.
Blends
One recipe below is ours; the other three are pairings to try in the widget. Sienna keeps the pronunciation in each.
Promo hook, an even split, the recipe we kept. Brooke is our question-and-answer pick; listen for a hook that sounds spoken rather than announced.
50% Brooke
A pairing to try. Our note on Delaney is a lower, throatier delivery; listen for whether half of it gives an ad read some edge without losing Sienna's speed.
40% to 50% Delaney
A pairing to try. Walter is our seasonal-announcement pick; listen for whether his weight helps a countdown.
50% Walter
A pairing to try for scripts longer than a hook. Maya's note is warm and even; listen for whether 60 percent of her steadies Sienna over two minutes.
60% Maya
Bright and quick, for openers and ads.
Sienna 50% + Brooke 50%
A fixed sample at this mix, rendered once and cached. No credits used.
Every blend works the same way: the first voice sets the pronunciation and pacing, the second voice lends only its timbre, and you choose the ratio between 10 and 90 percent. See the three steps on the voice library page
At a glance
| Language | English (US) |
|---|---|
| Voice | Female |
| Overall grade | A- |
| Target quality | A |
| Training audio | 10 to 100 hours |
| Captions | Word-level captions |
| Cost | 1 credit per job, room for 10,000 characters |
| Formats | MP3 free, WAV on paid plans |
| Commercial use | Included on every plan |
| Preview | Free, no sign-up |
Grade, target quality, and training audio come from the engine's published voice list and estimate training data, not how a voice sounds.
FAQ
Both are American English voices. Maya is graded A overall; Sienna is A- with a target quality of A and 10 to 100 hours of training audio. Our note calls Sienna the brighter of the two.
Start at 1.0 and listen. The published guidance says voices can rush on long utterances, so a shorter rewrite usually beats a faster setting.
Yes. The primary keeps its pronunciation and Sienna lends only her timbre. The widget below lets you try any share from 10 to 90 percent.
Yes. Every English voice reports word timing during synthesis, so the SRT and VTT files follow the read word by word.
Similar voices
Paste a script, keep Sienna or blend a second voice, and download the MP3 with subtitles.