Brief announcements
Notices, openings, and short spots, the use our note names. One sentence per line so every caption cue lands on a real pause.

Six recipes we kept start from Ruoxi, and in every one she holds only 10 to 20 percent of the mix. Her own record is modest, a D grade with C target quality and 10 to 100 minutes of training audio; our note gives her brief announcements.
Hear it
On the record
The kept recipes put her at 10 or 20 percent, yet she decides how each character is read: the first voice supplies pronunciation and pacing, the second only lends timbre. With Ruoxi in front, an English, British, or Japanese partner changes the colour of the sound and nothing about the Mandarin.
All eight Mandarin voices share the D grade, 10 to 100 minutes each and one to ten hours for the set, which is why we kept looking for partners.
Length the engine likes. The published guidance puts the sweet spot at roughly 100 to 200 tokens, a few sentences; very short lines are the weak spot and very long ones can rush.
Where we would use it
Notices, openings, and short spots, the use our note names. One sentence per line so every caption cue lands on a real pause.
Six kept recipes prove it: Ruoxi in front, any language behind her, the Mandarin intact.
Soft Mandarin pairs her with Tongtong, gentler than either alone by our note. A start for lesson material.
Blends
Every card is a blend we listened to and kept; the percentage is the second voice's share, and Ruoxi keeps the pronunciation.
Husky announcer: Mandarin announcements with a throaty edge, by our note. The most extreme ratio we kept.
80% Delaney
Clear Mandarin: Mandarin with the clarity of our grade A voice. Try it first if the plain read feels thin.
80% Maya
Soft Mandarin: two Mandarin voices, gentler than either alone by our note.
80% Tongtong
Mandarin, Japanese tone: the pronunciation is untouched; only the timbre is Japanese.
90% Yukiko
Mandarin announcements with a throaty edge.
Ruoxi 20% + Delaney 80%
Mandarin pronunciation, British lesson timbre.
Ruoxi 10% + Freya 90%
Mandarin with Maya's clarity.
Ruoxi 20% + Maya 80%
Mandarin with a British instructional tone.
Ruoxi 10% + Nell 90%
2 more in the generator
A fixed sample at this mix, rendered once and cached. No credits used.
A cross-language blend: Ruoxi keeps the pronunciation, the second voice only lends its timbre.
Every blend works the same way: the first voice sets the pronunciation and pacing, the second voice lends only its timbre, and you choose the ratio between 10 and 90 percent. See the three steps on the voice library page
At a glance
| Language | Chinese |
|---|---|
| Voice | Female |
| Overall grade | D |
| Target quality | C |
| Training audio | 10 to 100 minutes |
| Captions | Sentence-level captions |
| Cost | 1 credit per job, room for about 3,300 characters |
| Formats | MP3 free, WAV on paid plans |
| Commercial use | Included on every plan |
| Preview | Free, no sign-up |
Grade, target quality, and training audio come from the engine's published voice list and estimate training data, not how a voice sounds.
FAQ
Because the first voice already controls every pronunciation and pause; the second can take most of the mix without changing a syllable.
No. Mandarin voices report no per-word timing, so each cue is anchored to a real pause and cut at 16 characters a line.
Yes, directly. Multi-reading characters resolve more reliably in simplified text, so convert a long traditional script or use the Pronounce as tag.
Billing counts UTF-8 bytes, three per character, so a single credit stretches to about 3,300 characters; the exact cost shows before you confirm.
Similar voices
Paste a script, keep Ruoxi or blend a second voice, and download the MP3 with subtitles.