Software tutorials
Click-by-click guides, the use our note names. Write one action per line so each caption lands on one step.

On the published voice list Harper sits at C: B target quality, 10 to 100 minutes of training audio. Our note is specific: step-by-step software tutorials. She is not in any recipe we kept, and she has no listened blends, so what follows is the record, the note, and pairings to try.
Hear it
On the record
A B target quality means the reference audio was clean and its transcript matched; the 10 to 100 minute band is the smaller of the two common bands, which is what holds her at C rather than C+.
Our note names one job, software tutorials that go step by step. That points at scripts full of menu names, button labels, and numbered actions, which is also where the Pronounce as tag earns its keep.
Length the engine likes. The published guidance puts the sweet spot at roughly 100 to 200 tokens, a few sentences; very short lines are the weak spot and very long ones can rush.
Where we would use it
Click-by-click guides, the use our note names. Write one action per line so each caption lands on one step.
Feature tours and settings explainers where interface labels must be read exactly as written.
As a second voice under a demo or explainer primary, a pairing suggested on Miles's page and not yet listened to.
Blends
No recipe we kept includes Harper, so every pairing below is one to try in the widget. Harper keeps the pronunciation; the second voice lends its timbre.
Maya is our grade A narration voice, warm and even by our note; listen for whether 60 percent of her steadies a long tutorial.
60% Maya
Miles is our product-demonstration pick; listen for a demo tone under Harper's step-by-step read.
50% Miles
Our note on Delaney is low and throaty; listen for whether half of her softens a dry tutorial.
50% Delaney
Elena is our instructional-video pick, graded C+; listen for whether two teaching voices together read more evenly than either alone.
50% Elena
A fixed sample at this mix, rendered once and cached. No credits used.
Every blend works the same way: the first voice sets the pronunciation and pacing, the second voice lends only its timbre, and you choose the ratio between 10 and 90 percent. See the three steps on the voice library page
At a glance
| Language | English (US) |
|---|---|
| Voice | Female |
| Overall grade | C |
| Target quality | B |
| Training audio | 10 to 100 minutes |
| Captions | Word-level captions |
| Cost | 1 credit per job, room for 10,000 characters |
| Formats | MP3 free, WAV on paid plans |
| Commercial use | Included on every plan |
| Preview | Free, no sign-up |
Grade, target quality, and training audio come from the engine's published voice list and estimate training data, not how a voice sounds.
FAQ
It is the engine's published estimate of training data: B target quality with only 10 to 100 minutes of audio. Voices with an hour or more of similar audio grade C+.
Wrap the label in the Pronounce as tag and type how it should sound. The label stays as written on screen and in the subtitles.
Yes. Put one action per line and add a pause tag where the cursor moves; the SRT that comes back shows where each line starts.
Yes. Standard voice output carries commercial rights on every plan, with no attribution and no watermark.
Similar voices
Paste a script, keep Harper or blend a second voice, and download the MP3 with subtitles.