Instructional videos
Setup guides and how-to clips, the use our note names. One step per line, a pause tag where the viewer needs time.

Instructional videos are the use our note gives Michiko. The published list records her training source, a reading of 手袋を買いに, Buying Mittens, Niimi Nankichi's story of a young fox sent to town, 10 to 100 minutes of it at B target quality, which places her at C.
Hear it
On the record
The list names her source and its length: a children's story, read aloud, tens of minutes. The grade follows the amount. Our note goes a different direction and gives her instructional videos, where the script is steps, menu names, and numbers.
Mixed scripts are the practical point: app names and buttons in Latin letters are read in place with a Japanese accent, numbers are read out in Japanese, and each step on its own line becomes one subtitle cue.
Length the engine likes. The published guidance puts the sweet spot at roughly 100 to 200 tokens, a few sentences; very short lines are the weak spot and very long ones can rush.
Where we would use it
Setup guides and how-to clips, the use our note names. One step per line, a pause tag where the viewer needs time.
Menu names in Latin letters are read in place by her, with a Japanese accent, so nothing is skipped.
Sentence-level SRT and WebVTT come with every job, cut at 16 characters a line.
Blends
Nothing here is a kept recipe. Each card is a pairing to try; Michiko keeps the readings and the second voice lends its timbre.
A pairing to try, two Japanese voices. Sakura's note is clean readings; listen for whether the pairing keeps them.
50% Sakura
A pairing to try. Harper is our English software-tutorial voice; listen for a tutorial timbre carried into Japanese.
50% Harper
A pairing to try. Delaney is the husky partner in both kept Japanese recipes; listen for whether it suits instructions.
50% Delaney
A pairing to try, the Japanese male with a slower cadence by our note; listen for whether he slows an instruction read.
40% Takeshi
A fixed sample at this mix, rendered once and cached. No credits used.
Every blend works the same way: the first voice sets the pronunciation and pacing, the second voice lends only its timbre, and you choose the ratio between 10 and 90 percent. See the three steps on the voice library page
At a glance
| Language | Japanese |
|---|---|
| Voice | Female |
| Overall grade | C |
| Target quality | B |
| Training audio | 10 to 100 minutes |
| Captions | Sentence-level captions |
| Cost | 1 credit per job, room for about 3,300 characters |
| Formats | MP3 free, WAV on paid plans |
| Commercial use | Included on every plan |
| Preview | Free, no sign-up |
Grade, target quality, and training audio come from the engine's published voice list and estimate training data, not how a voice sounds.
FAQ
Put one action per line and add a pause tag where the cursor moves. The SRT has one cue per line, anchored to the pause the engine made.
In place, by Michiko, with a Japanese accent, the way a presenter would say them. Tag any name that should sound different with Pronounce as.
Japanese voices report no per-word timing, so cues are anchored to real pauses and capped at 16 characters a line.
It is the source recording named on the published list: a reading of Niimi Nankichi's story, 10 to 100 minutes. A fact about the data, not the tone.
Similar voices
Paste a script, keep Michiko or blend a second voice, and download the MP3 with subtitles.