Sakura, Japanese voice

Sakura, the Japanese voice with the most training audio on the list

Sakura holds the best record of the five Japanese voices: a C+ grade, B target quality, and 1 to 10 hours of training audio where the others have minutes. Our note is about accuracy, a clean reading of kanji and kana, and two recipes we kept start from her, both with Delaney.

JapaneseFemaleC+Sentence-level captionsStandard voice

Hear it

Three samples, three registers

  1. Listen for the readings on 港 and 目, the kind of kanji our note says she gets clean.
  2. Listen for 三時間 in the middle of the hook, a number read the Japanese way rather than as digits.
  3. Listen for the break between the two sentences; the caption cue is cut at that pause.

On the record

An hour or more of audio, and what that buys

Four of the five Japanese voices were trained on minutes of audio; Sakura has hours, which is the whole difference between her C+ and their C and C-. Target quality B means the reference recordings were reasonably clean.

Our note, clean reading of kanji and kana, is about readings, not warmth. It is why both Japanese recipes we kept start from her: the second voice changes the timbre while her readings stay put.

Length the engine likes. The published guidance puts the sweet spot at roughly 100 to 200 tokens, a few sentences; very short lines are the weak spot and very long ones can rush.

Where we would use it

Narration where readings matter, and blends that keep them

Narration and voiceover

Scripts heavy with kanji, the use our note points at. Readings come from context; tag the rare ones with Pronounce as.

Generate Japanese with Sakura

Mixed Japanese and English

Product names in Latin letters are read in place by her, with a Japanese accent.

Open the main generator

The base of two kept recipes

Husky Japanese and its lighter version both start from her. Try them before building your own.

Browse every voice

Blends

Two kept recipes with Delaney, one dial between them

The Delaney card covers two recipes we listened to and kept; the others are pairings to try. Sakura reads every character in each.

Delaney, English (US) voice
DelaneyEnglish (US)

With Delaney

Husky Japanese at 80 percent Delaney, and a lighter version at an even split that keeps her clean kana, by our note.

80% or 50% Delaney

Maya, English (US) voice
MayaEnglish (US)

With Maya

A pairing to try. Our grade A voice at 60 percent behind Sakura; listen for whether the clarity carries across languages.

60% Maya

Yukiko, Japanese voice
YukikoJapanese

With Yukiko

A pairing to try, two Japanese voices: Yukiko is our story-excerpt voice by our note.

50% Yukiko

Takeshi, Japanese voice
TakeshiJapanese

With Takeshi

A pairing to try, the only Japanese male on the list; our note gives him a slower cadence.

50% Takeshi

Recipes we kept with Sakura

Husky Japanese

Japanese reading, mostly husky timbre.

Sakura 20% + Delaney 80%

Husky Japanese · lighter

Even split, keeping Sakura's clean kana.

Sakura 50% + Delaney 50%

Blend Sakura with any second voice

A fixed sample at this mix, rendered once and cached. No credits used.

A cross-language blend: Sakura keeps the pronunciation, the second voice only lends its timbre.

Every blend works the same way: the first voice sets the pronunciation and pacing, the second voice lends only its timbre, and you choose the ratio between 10 and 90 percent. See the three steps on the voice library page

At a glance

Sakura on the published list

LanguageJapanese
VoiceFemale
Overall gradeC+
Target qualityB
Training audio1 to 10 hours
CaptionsSentence-level captions
Cost1 credit per job, room for about 3,300 characters
FormatsMP3 free, WAV on paid plans
Commercial useIncluded on every plan
PreviewFree, no sign-up

Grade, target quality, and training audio come from the engine's published voice list and estimate training data, not how a voice sounds.

FAQ

Questions about Sakura

  • Why is Sakura graded higher than the other Japanese voices?

    Training audio: 1 to 10 hours where the others have minutes. The grade tracks that amount, not how she sounds.

  • What happens to English words in a Japanese script?

    Read in place by Sakura with a Japanese accent, as a presenter would say a product name. For an English-first script, choose an English voice.

  • How do I correct a kanji reading?

    Wrap the word in the Pronounce as tag and type the kana. The kanji stays on screen and in the subtitles; only the voice follows the kana.

  • Are the subtitles timed per word?

    Per sentence. Japanese voices report no word timing, so each cue is anchored to a real pause and cut at 16 characters a line.

Similar voices

Other Japanese voices to compare

Generate with Sakura

Paste a script, keep Sakura or blend a second voice, and download the MP3 with subtitles.

Browse all voices