Short how-to videos
One task from start to finish, the use our note names. One step per line.

Caleb carries a D on the published list, C target quality with 10 to 100 minutes of training audio. Our note gives him short how-to videos: one task, one or two minutes. No recipe we kept includes him, so the pairings below are suggestions to hear in the widget.
Hear it
On the record
The D grade follows the published rule: C-quality reference audio, and only minutes of it. It shares that record with Dean, Miles, and Marcus among the American males.
Our note is about format rather than tone: short how-to videos. Keeping each utterance to a few sentences also matches the published guidance, which names very short and very long lines as the weak spots.
Length the engine likes. The published guidance puts the sweet spot at roughly 100 to 200 tokens, a few sentences; very short lines are the weak spot and very long ones can rush.
Where we would use it
One task from start to finish, the use our note names. One step per line.
Vertical clips that show one thing. Keep the script under a minute.
Untested as a second voice; the widget lets you hear him under any primary.
Blends
None of these has been listened to. Caleb keeps the pronunciation; the second voice lends its timbre.
Brooke is our question-and-answer pick; listen for a more conversational how-to.
50% Brooke
Miles is our product-demonstration pick; listen for whether two demo voices read a task more steadily.
50% Miles
Our note on Delaney is low and throaty; listen for whether half of her softens the read.
50% Delaney
Maya is warm and even by our note; listen for whether 60 percent of her lifts a D-graded voice.
60% Maya
A fixed sample at this mix, rendered once and cached. No credits used.
Every blend works the same way: the first voice sets the pronunciation and pacing, the second voice lends only its timbre, and you choose the ratio between 10 and 90 percent. See the three steps on the voice library page
At a glance
| Language | English (US) |
|---|---|
| Voice | Male |
| Overall grade | D |
| Target quality | C |
| Training audio | 10 to 100 minutes |
| Captions | Word-level captions |
| Cost | 1 credit per job, room for 10,000 characters |
| Formats | MP3 free, WAV on paid plans |
| Commercial use | Included on every plan |
| Preview | Free, no sign-up |
Grade, target quality, and training audio come from the engine's published voice list and estimate training data, not how a voice sounds.
FAQ
It is the published estimate of his training data: C-quality audio measured in minutes. The list says sound itself is subjective.
Write one action per line and cut anything the footage already shows. The generator shows the byte count and credit cost before you submit.
Yes. English voices report word timing during synthesis, so each SRT cue starts on the spoken word.
Yes. Standard voice output is cleared for commercial use on every plan, with no watermark or credit line.
Similar voices
Paste a script, keep Caleb or blend a second voice, and download the MP3 with subtitles.