Tutorial steps
Click-by-click guides, the use our note names. One action per line, a pause tag where the cursor moves.

Brief tutorial steps are the use our note gives Jiahao, and he has no kept recipe yet, so this page leans on the record and on the mechanics of a step-by-step script. The record is the standard Mandarin one: D grade, C target quality, 10 to 100 minutes of training audio, part of a Mandarin set that totals one to ten hours.
Hear it
On the record
A tutorial step in Mandarin usually contains a menu name, a number, and one verb. Latin menu names are read as English automatically; numbers read cleanly; the verb is where multi-reading characters bite, so Tidy up script is worth a click before generating.
Sentence-level captions suit tutorials well: one step per line gives one cue per step, anchored to the pause after the full stop, with lines capped at 16 characters so they fit a phone screen.
Length the engine likes. The published guidance puts the sweet spot at roughly 100 to 200 tokens, a few sentences; very short lines are the weak spot and very long ones can rush.
Where we would use it
Click-by-click guides, the use our note names. One action per line, a pause tag where the cursor moves.
Two or three steps to solve one problem, short enough that a level read is all that is needed.
Not yet listened to in any blend; the pairings below are our first candidates.
Blends
No kept recipe includes Jiahao. Each card is a pairing to try; he keeps the pronunciation and the second voice lends its timbre.
A pairing to try, two Mandarin voices with instruction notes: Yuwen reads short instructions by our note.
50% Yuwen
A pairing to try. Our grade A voice at 60 percent behind a tutorial read; listen for whether she steadies the steps.
60% Maya
A pairing to try. Harper is our English software-tutorial voice; listen for a tutorial timbre carried across languages.
50% Harper
A pairing to try, two Mandarin males. Weilun has five kept recipes, so the timbre he lends is documented.
50% Weilun
A fixed sample at this mix, rendered once and cached. No credits used.
Every blend works the same way: the first voice sets the pronunciation and pacing, the second voice lends only its timbre, and you choose the ratio between 10 and 90 percent. See the three steps on the voice library page
At a glance
| Language | Chinese |
|---|---|
| Voice | Male |
| Overall grade | D |
| Target quality | C |
| Training audio | 10 to 100 minutes |
| Captions | Sentence-level captions |
| Cost | 1 credit per job, room for about 3,300 characters |
| Formats | MP3 free, WAV on paid plans |
| Commercial use | Included on every plan |
| Preview | Free, no sign-up |
Grade, target quality, and training audio come from the engine's published voice list and estimate training data, not how a voice sounds.
FAQ
Put one action per line and add a pause tag where the cursor moves. The SRT that comes back has one cue per line, anchored to the pause the engine made, so it lines up with the steps.
They are split from the Mandarin by writing system and read as English, with no markup needed. Tag anything that should sound different with Pronounce as.
We only publish blends we listened to and kept, and none starts from Jiahao yet. The widget lets you preview any pairing for free.
It is the published estimate of how much audio the voice was trained on, the smaller of the two common bands. It explains the D grade; the samples show what the voice actually does.
Similar voices
Paste a script, keep Jiahao or blend a second voice, and download the MP3 with subtitles.