Short instructions
Step lists, setup guides, and safety notes, the use our note names. One step per line, a pause tag before each action.

Our note on Yuwen is one line, short Mandarin instructions, and no recipe we kept includes her, so this page has less listening behind it than Ruoxi's or Weilun's. What it has is the record, a D grade with C target quality and 10 to 100 minutes of training audio, and the mechanics that matter for instructions: pauses, readings, and captions.
Hear it
On the record
An instruction script lives on its pauses. The engine gives about four tenths of a second at a comma or full stop, and a pause tag inserts more where a viewer needs time to act. Captions are cut at those same pauses, so one step per line gives one cue per step.
Readings matter more here than in narration: 行 as a table row, 长 as length, 重 as again. The engine takes the most frequent reading, so run Tidy up script or add the Pronounce as tag for any step that depends on the right one.
Length the engine likes. The published guidance puts the sweet spot at roughly 100 to 200 tokens, a few sentences; very short lines are the weak spot and very long ones can rush.
Where we would use it
Step lists, setup guides, and safety notes, the use our note names. One step per line, a pause tag before each action.
Menu names and buttons read exactly as written. English labels inside the Mandarin are read as English automatically.
Not yet listened to in any blend; the pairings below are where we would start.
Blends
Nothing here is a kept recipe. Each card is a pairing we would try first; Yuwen keeps the pronunciation and the second voice lends timbre.
A pairing to try, two Mandarin females. Ruoxi is the base of six kept recipes, so her timbre is the best documented in the set.
50% Ruoxi
A pairing to try. Clear Mandarin with Ruoxi kept Maya at 80 percent; listen for whether the same share lifts an instruction read.
80% Maya
A pairing to try. Instructional Mandarin with Ruoxi kept Nell at 90 percent; Yuwen in front of the same partner is the obvious next test.
90% Nell
A pairing to try, two Mandarin voices with tutorial notes: Jiahao reads brief tutorial steps by our note.
50% Jiahao
A fixed sample at this mix, rendered once and cached. No credits used.
Every blend works the same way: the first voice sets the pronunciation and pacing, the second voice lends only its timbre, and you choose the ratio between 10 and 90 percent. See the three steps on the voice library page
At a glance
| Language | Chinese |
|---|---|
| Voice | Female |
| Overall grade | D |
| Target quality | C |
| Training audio | 10 to 100 minutes |
| Captions | Sentence-level captions |
| Cost | 1 credit per job, room for about 3,300 characters |
| Formats | MP3 free, WAV on paid plans |
| Commercial use | Included on every plan |
| Preview | Free, no sign-up |
Grade, target quality, and training audio come from the engine's published voice list and estimate training data, not how a voice sounds.
FAQ
Insert a pause tag at the start of the line, or end the previous line with a full stop for the engine's own pause. Longer pauses are cut into the audio as real silence.
One step per line gives one cue per step, anchored to the pause between them and cut at 16 characters a line. Mandarin voices report no per-word timing, so the cues are sentence-level.
Yes. Text is split by writing system before synthesis, so app names and button labels in Latin letters are read as English inside the Mandarin.
We only list blends we listened to and kept, and none starts from Yuwen yet. The widget lets you try the pairings above, free and without signing in.
Similar voices
Paste a script, keep Yuwen or blend a second voice, and download the MP3 with subtitles.