Narration and voiceover
Scripts heavy with kanji, the use our note points at. Readings come from context; tag the rare ones with Pronounce as.

Sakura holds the best record of the five Japanese voices: a C+ grade, B target quality, and 1 to 10 hours of training audio where the others have minutes. Our note is about accuracy, a clean reading of kanji and kana, and two recipes we kept start from her, both with Delaney.
Hear it
On the record
Four of the five Japanese voices were trained on minutes of audio; Sakura has hours, which is the whole difference between her C+ and their C and C-. Target quality B means the reference recordings were reasonably clean.
Our note, clean reading of kanji and kana, is about readings, not warmth. It is why both Japanese recipes we kept start from her: the second voice changes the timbre while her readings stay put.
Length the engine likes. The published guidance puts the sweet spot at roughly 100 to 200 tokens, a few sentences; very short lines are the weak spot and very long ones can rush.
Where we would use it
Scripts heavy with kanji, the use our note points at. Readings come from context; tag the rare ones with Pronounce as.
Product names in Latin letters are read in place by her, with a Japanese accent.
Husky Japanese and its lighter version both start from her. Try them before building your own.
Blends
The Delaney card covers two recipes we listened to and kept; the others are pairings to try. Sakura reads every character in each.
Husky Japanese at 80 percent Delaney, and a lighter version at an even split that keeps her clean kana, by our note.
80% or 50% Delaney
A pairing to try. Our grade A voice at 60 percent behind Sakura; listen for whether the clarity carries across languages.
60% Maya
A pairing to try, two Japanese voices: Yukiko is our story-excerpt voice by our note.
50% Yukiko
A pairing to try, the only Japanese male on the list; our note gives him a slower cadence.
50% Takeshi
Japanese reading, mostly husky timbre.
Sakura 20% + Delaney 80%
Even split, keeping Sakura's clean kana.
Sakura 50% + Delaney 50%
A fixed sample at this mix, rendered once and cached. No credits used.
A cross-language blend: Sakura keeps the pronunciation, the second voice only lends its timbre.
Every blend works the same way: the first voice sets the pronunciation and pacing, the second voice lends only its timbre, and you choose the ratio between 10 and 90 percent. See the three steps on the voice library page
At a glance
| Language | Japanese |
|---|---|
| Voice | Female |
| Overall grade | C+ |
| Target quality | B |
| Training audio | 1 to 10 hours |
| Captions | Sentence-level captions |
| Cost | 1 credit per job, room for about 3,300 characters |
| Formats | MP3 free, WAV on paid plans |
| Commercial use | Included on every plan |
| Preview | Free, no sign-up |
Grade, target quality, and training audio come from the engine's published voice list and estimate training data, not how a voice sounds.
FAQ
Training audio: 1 to 10 hours where the others have minutes. The grade tracks that amount, not how she sounds.
Read in place by Sakura with a Japanese accent, as a presenter would say a product name. For an English-first script, choose an English voice.
Wrap the word in the Pronounce as tag and type the kana. The kanji stays on screen and in the subtitles; only the voice follows the kana.
Per sentence. Japanese voices report no word timing, so each cue is anchored to a real pause and cut at 16 characters a line.
Similar voices
Paste a script, keep Sakura or blend a second voice, and download the MP3 with subtitles.