Japanese TTS tool

Japanese Text to Speech for Kanji, Kana, and Mixed-in English

Paste Japanese text exactly as you write it, choose a Japanese voice, and download an MP3 with sentence-timed subtitles at 16 characters a line. The Japanese TTS is free to use, with no watermark and no sign-up to listen.

  • 5Japanese voices
  • 16characters per subtitle line
  • 5free credits every month
0 bytes / 0 jobs / 0 credits

You can edit or delete these tags directly in your script.

Voice
SakuraJapanese · Female
Advanced settings

Output

0.5x1x2x
Output format

Voice mix

The first voice sets pronunciation; the second only lends its character.

Primary voice

Sakura

Sets pronunciation and pacing.

Second voice

A fixed sample at your current mix and speed. No credits used.
  • Commercial use included on the free plan
  • No watermark on any output
  • Preview every voice without signing up

Voices for this page

Japanese TTS voices, female and male

Four female voices and one male voice, all grade C+ to C- on the engine's published list, which grades training data rather than sound. Our notes point them at lessons, walkthroughs, and story excerpts. Play the samples before you spend a credit.

What you get

What this Japanese voice generator handles

Japanese scripts mix three writing systems and, increasingly, English product names. Each of those is dealt with, and so are the readings.

Kanji, kana, and numbers read in context

The engine chooses readings from the surrounding sentence, and dates, prices, and other numbers are read out in Japanese rather than as digits.

When a reading is wrong

Homographs are the one place the engine guesses. It always takes the most common reading, so a rarer one such as なまもの for 生物 needs a hint. Wrap the word in the Pronounce as tag and type the kana you want; the script keeps the kanji on screen and the voice reads your kana. Tidy up script can add these tags for you.

Listen to every Japanese voice

English words inside Japanese text

Product names, app menus, and loanwords written in the Latin alphabet are read in place by the Japanese voice, with a Japanese accent, instead of being skipped.

How mixed scripts are rendered

The script is split by writing system: Japanese runs go through the Japanese reader, Latin runs through the English reader, and the two are joined into one take by the same voice. The accent on the English is part of the character, the way a Japanese presenter would say it. For an English-first script, an English voice is the better choice.

Working in Mandarin as well? See the Chinese page

Pace for listeners and learners

Set the speed anywhere from 0.5 to 2 times. Half speed keeps the pitch natural, which makes it usable for shadowing and dictation practice.

Pause and emphasis tags

Japanese punctuation already produces natural pauses. For a longer gap, a pause tag sets the exact seconds of silence, and a slow tag reads a single phrase at a reduced rate for emphasis. Both stay visible in the script and can be removed with one click, and neither costs extra credits.

Need another language? Open the main generator

Subtitles

Japanese subtitles at 16 characters a line

Japanese subtitles need short lines and clean breaks at the particles. The SRT here is cut for that, and it comes out of the same render as the audio.

  • SRT and WebVTT with every job, on the free plan included
  • Lines capped at 16 characters, the usual width for Japanese captions
  • Cues anchored to real pauses in the audio, aligned per sentence
  • Edit a sentence, render again, and the cues move with the audio

The honest limit: Japanese voices report timing per sentence, not per word. The worker finds the pauses in the finished audio and anchors each cue to them, which is accurate enough for normal captions but not for karaoke-style word highlighting. The English voices on this site do return word-level timing.

Generate with subtitles
1
00:00:00,180 --> 00:00:03,240
今日は新しい機能を紹介します。

2
00:00:03,520 --> 00:00:06,900
設定画面を開いてください。

3
00:00:07,180 --> 00:00:10,400
右上のボタンを押すだけです。

Voice blends

Two Japanese voice blends with a different timbre

Both blends keep Sakura's clean kana reading and add Delaney's huskier character at two strengths. The first voice sets the pronunciation, so the Japanese stays correct.

Husky Japanese

Japanese reading, mostly husky timbre.

Sakura 20% + Delaney 80%

Husky Japanese · lighter

Even split, keeping Sakura's clean kana.

Sakura 50% + Delaney 50%

How voice blending works

The primary voice decides how every word is read and paced; the second voice only lends its timbre, at the percentage you set. The second voice can be from any language, which is how a Japanese reading gains an English voice's warmth without changing a syllable. Preview a blend at no cost, then save it in this browser.

How it works

How to convert Japanese text to speech, step by step

  1. A paper lantern with sound waves on either side, standing in for Japanese text read aloud.STEP 01

    Paste the Japanese text

    Kanji, hiragana, katakana, and any English in between all go in as they are. The generator shows the byte count and the exact credit cost before you submit.

  2. Two tuning forks of different sizes, one for each voice to compare.STEP 02

    Pick a Japanese voice

    Play the samples on the five cards, choose a voice, and set the speed. Blend in a second voice from the advanced settings if you want more warmth.

  3. A magnifying glass over a page of text, for checking readings before generating.STEP 03

    Check the readings, then download

    Run Tidy up script or add a Pronounce as tag where you know the engine will miss. The MP3 and the SRT and VTT files then appear in My Speech.

Free plan, plainly

Free Japanese TTS, and what a paid plan adds

Everything on this page, including subtitles and commercial use, works on the free plan.

CapabilityFreeStarterCreator
PriceFree$6.99 a month, or $69.99 a year$15.99 a month, or $159.99 a year
Credits per month5 credits100 credits300 credits
Standard English characters per job10,000 for 1 credit10,000 for 1 credit10,000 for 1 credit
QueueFree queue, no promised timeSkips the free queueSkips the free queue
Output formatMP3MP3 and WAVMP3 and WAV
SubtitlesSRT and VTTSRT and VTTSRT and VTT
Script tidy (AI)Not availableIncludedIncluded
Premium voicesNot availableComing soonComing soon
Credit packsNot available100 credits for $5100 credits for $5
File retention72 hours30 days30 days
Commercial useIncludedIncludedIncluded
WatermarkNoneNoneNone
How credits are counted

One credit runs one job, and that job comes with room for 10,000 UTF-8 bytes: about 10,000 characters of English, or about 3,300 Chinese or Japanese characters. That is a whole article or a long script in a single go. Longer pieces are never turned away either, since each further block of 10,000 bytes adds a credit and what comes back is still one continuous file. Credits are reserved when you submit and refunded in full if a job fails.

Compare the plans

FAQ

Questions about Japanese TTS, answered

  • Does it read kanji correctly?

    Mostly, yes. Readings are chosen from context, and common words are reliable. Homographs are the weak spot: the engine picks the most frequent reading, so a rarer one needs a hint from you.

  • Can I fix a wrong reading?

    Yes. Wrap the word in the Pronounce as tag and type the kana; the kanji stays on screen and the voice reads your kana. Tidy up script scans for homographs and adds the tag where it is confident.

  • Are there male Japanese voices?

    One, Takeshi, alongside four female voices. All five sit at grade C+ to C- on the published list, which reflects how much audio each was trained on; listen to the samples to hear the differences in tone.

  • Can I slow the speech down for language practice?

    Yes. The speed slider goes down to 0.5 times with the pitch unchanged, and the sentence-level subtitles double as a script to read along with.

  • Does it support furigana?

    Not as ruby text. Paste plain Japanese and give a reading with the Pronounce as tag where needed; for the voice, the tag does the same job that furigana does for a reader.

  • Why does Japanese text use credits faster than English?

    Standard billing counts UTF-8 bytes, and each kana or kanji is three bytes, so a single credit carries about 3,300 Japanese characters in one job. Anything longer runs in one pass too, one further credit per block. The generator shows the exact credit cost before you click generate.

  • Is the audio free to use in my videos?

    Yes, on every plan. Standard voice output is licensed for commercial use with no watermark and no attribution. Subtitles come as a separate SRT or VTT file with sentence-level alignment.

Generate your Japanese audio

Paste the text, pick a voice, and download the MP3 with Japanese subtitles.