Lesson examples
Single sentences for learners to repeat, the use our note names. One example per line gives one caption per example.

Tongtong appears in one recipe we kept, and not in front: Soft Mandarin puts her at 80 percent behind Ruoxi, and our note calls the result gentler than either voice alone. Her own assignment is short lesson examples. On the published list she is a D with C target quality and 10 to 100 minutes of training audio.
Hear it
On the record
Soft Mandarin is unusual among our recipes because both voices are Mandarin. Ruoxi in front sets the pronunciation; Tongtong at 80 percent supplies most of the timbre. Our listening note is short: gentler than either alone.
Lesson examples are short by nature, a sentence read once so a learner can repeat it, and sentence-level captions cut at real pauses fit that pattern. Her record matches every Mandarin voice here.
Length the engine likes. The published guidance puts the sweet spot at roughly 100 to 200 tokens, a few sentences; very short lines are the weak spot and very long ones can rush.
Where we would use it
Single sentences for learners to repeat, the use our note names. One example per line gives one caption per example.
Word lists and short phrases where a level, unhurried read matters more than expression.
The kept recipe with Ruoxi in front, gentler than either alone by our note. Try it for longer lesson passages.
Blends
Soft Mandarin is the recipe we kept, with Tongtong as the second voice. The cards below put her in front instead; each is a pairing to try except where noted.
Reverse of Soft Mandarin: Tongtong in front, Ruoxi behind. The kept version runs the other way, so treat this as a pairing to try.
50% Ruoxi
A pairing to try for lessons. Nora is our lesson-introduction voice; listen for whether 60 percent of her gives an example sentence a teacher's frame.
60% Nora
A pairing to try. Imogen is the top-graded British voice on the list at B-, with a courseware note; listen for a courseware timbre on Mandarin.
60% Imogen
A pairing to try. Delaney supplies the throaty edge in most kept Mandarin blends; listen for whether it suits a lesson.
40% Delaney
Two Mandarin voices blended, gentler than either alone.
Ruoxi 20% + Tongtong 80%
A fixed sample at this mix, rendered once and cached. No credits used.
Every blend works the same way: the first voice sets the pronunciation and pacing, the second voice lends only its timbre, and you choose the ratio between 10 and 90 percent. See the three steps on the voice library page
At a glance
| Language | Chinese |
|---|---|
| Voice | Female |
| Overall grade | D |
| Target quality | C |
| Training audio | 10 to 100 minutes |
| Captions | Sentence-level captions |
| Cost | 1 credit per job, room for about 3,300 characters |
| Formats | MP3 free, WAV on paid plans |
| Commercial use | Included on every plan |
| Preview | Free, no sign-up |
Grade, target quality, and training audio come from the engine's published voice list and estimate training data, not how a voice sounds.
FAQ
She is 80 percent of it, with Ruoxi in front for pronunciation. Our note calls the blend gentler than either alone.
Yes. Put each example on its own line and end it with a full stop; the caption cue is anchored to the pause the engine makes there, with lines capped at 16 characters.
Latin text is split off and read as English, so a pinyin string will be read as English letters, not as Mandarin. Write the characters instead and let the voice pronounce them.
They estimate training data, not sound: 10 to 100 minutes of audio at C target quality. The samples above are the only way to judge the voice.
Similar voices
Paste a script, keep Tongtong or blend a second voice, and download the MP3 with subtitles.