Long-form narration tool

Narrator Text to Speech for Long Scripts, in One Take

Paste a full article, script, or book excerpt and get one continuous narration back, with a subtitle file timed to it. Calm narrative voices, no fixed cap on a single job, and free to start.

  • 31narrative English voices
  • 10,000characters per job, one credit
  • 5free credits every month
0 bytes / 0 jobs / 0 credits

You can edit or delete these tags directly in your script.

Voice
MayaEnglish (US) · Female
Advanced settings

Output

0.5x1x2x
Output format

Voice mix

The first voice sets pronunciation; the second only lends its character.

Primary voice

Maya

Sets pronunciation and pacing.

Second voice

A fixed sample at your current mix and speed. No credits used.
  • Commercial use included on the free plan
  • No watermark on any output
  • Preview every voice without signing up

Voices for this page

Narrative voices built for long reads

Narration and documentary voices are listed first. Each card plays a calm narration sample, so you can judge whether a voice holds steady over a paragraph rather than just a line.

What you get

What a narrator voice generator needs to handle

Length, pacing, and consistency are the whole job in long-form narration. Here is how each one is handled.

The whole script in one job

There is no per-job character cap. A 30,000-character script goes through as one render and comes back as a single MP3 with one continuous subtitle track.

How long jobs are rendered

The engine works through the script in blocks and stitches them into one file, so the voice and the pace stay identical from the first line to the last. One credit carries 10,000 characters of English, so a full piece goes in as one submission for a handful of credits; a Starter month is 100 of them. Paid renders stay available for 30 days.

Compare the narration samples on every voice

Pauses and emphasis you control

Insert a pause of up to five seconds anywhere, slow a phrase that carries weight, or let Tidy up script place the pauses for you.

What the tags do

Blank lines and sentence endings already produce natural pauses. For anything longer, a pause tag sets the exact seconds of silence. A slow tag reads one phrase at a reduced rate, which is how a narrator lands a turning point without raising the voice. A pronunciation tag fixes a name the engine gets wrong, and every tag stays visible in the script.

Cutting something short instead? Try the TikTok page

One voice, start to finish

Set the speed once, between 0.5 and 2 times, and it holds across the whole render. Save the voice and its settings to reuse on the next piece.

Keeping a series consistent

Up to eight voices, blends and speed included, can be saved in this browser, and your last chosen voice is remembered on your account. Standard voices carry an honest quality grade from A to D. Maya, graded A, is the safest first pick for narration; play the calm sample on any card to compare.

Narrating in another language? Open the main generator

Subtitles

Subtitles that stay in sync across a long narration

A subtitle file generated separately from the audio drifts a little more with every minute. Here both come from the same render, so minute twenty lines up as well as minute one.

  • SRT and WebVTT returned with every narration, free plan included
  • Word-level timing on standard English voices, measured during synthesis
  • Cues break at real pauses, capped at 42 characters a line
  • Edit a paragraph, render again, and the cues follow

For a video, drop the MP3 and the SRT into the timeline at the same start point and leave them alone; the sync holds to the end. For an audiobook, the timed file doubles as a transcript, which makes it easy to locate and re-record a single sentence later without touching the rest.

Narrate with subtitles
1
00:00:00,140 --> 00:00:04,620
The house had been empty for eleven years,

2
00:00:04,700 --> 00:00:08,900
long enough for the garden to forget it.

3
00:00:09,260 --> 00:00:13,480
Nobody in the village said so out loud.

Voice blends

Audiobook and documentary blends, ready to use

These blends pair a clean primary voice with a warmer or deeper second voice. The first voice keeps the pronunciation; the second adds body for a long listen.

Audiobook narrator

Even and warm with husk underneath; holds a long read.

Maya 50% + Delaney 50%

Warm narrator

Male warmth with Maya's clean pronunciation.

Maya 30% + Walter 70%

Documentary voice

Deep male with a throaty edge.

Grant 50% + Delaney 50%

Neutral duo

Half clear female, half neutral male.

Maya 50% + Grant 50%

Course intro

Teacherly opening with a touch of warmth.

Delaney 30% + Nora 70%

Product demo

Steady male demo tone with husk underneath.

Delaney 30% + Miles 70%

How voice blending works

The primary voice decides how every word is pronounced and paced; the second voice contributes timbre only, at the percentage you set. The generator renders a fixed sample at your current mix and speed at no cost, so you can compare a blend against the plain voice before committing a long script to it. Save the ones that work.

How it works

How to narrate long text with AI, start to finish

  1. A long scroll of ruled text unrolling downward, standing in for a full script.STEP 01

    Paste the full text

    Headings, paragraphs, and blank lines all survive the paste and shape the pauses. The generator shows the byte count and the exact credit cost before you submit.

  2. Over-ear headphones resting on a closed book, for choosing a narrator voice.STEP 02

    Choose a narrative voice and a pace

    Play the narration sample on a few cards, pick one, and set the speed. Around 0.9 times suits audiobooks; 1.0 suits documentaries and explainers.

  3. A strip of film with caption lines under each frame, for the subtitle download.STEP 03

    Download one file, plus the subtitles

    The finished MP3 and its SRT and VTT files appear in My Speech. Long jobs keep rendering if you close the tab, so come back when it is done.

Free plan, plainly

Free narration, and what a paid plan adds

Long scripts are where credits matter. Here is what each plan gives you.

CapabilityFreeStarterCreator
PriceFree$6.99 a month, or $69.99 a year$15.99 a month, or $159.99 a year
Credits per month5 credits100 credits300 credits
Standard English characters per job10,000 for 1 credit10,000 for 1 credit10,000 for 1 credit
QueueFree queue, no promised timeSkips the free queueSkips the free queue
Output formatMP3MP3 and WAVMP3 and WAV
SubtitlesSRT and VTTSRT and VTTSRT and VTT
Script tidy (AI)Not availableIncludedIncluded
Premium voicesNot availableComing soonComing soon
Credit packsNot available100 credits for $5100 credits for $5
File retention72 hours30 days30 days
Commercial useIncludedIncludedIncluded
WatermarkNoneNoneNone
How credits are counted

One credit runs one job, and that job comes with room for 10,000 UTF-8 bytes: about 10,000 characters of English, or about 3,300 Chinese or Japanese characters. That is a whole article or a long script in a single go. Longer pieces are never turned away either, since each further block of 10,000 bytes adds a credit and what comes back is still one continuous file. Credits are reserved when you submit and refunded in full if a job fails.

See the plans side by side

FAQ

Narrator TTS questions, answered

  • How long a text can I narrate in one job?

    There is no fixed cap. One credit carries 10,000 characters of English, so a 45,000-character script goes in as a single submission for five credits. The result is one continuous MP3, not a set of pieces.

  • Which voice is best for audiobook narration?

    Start with Maya, the only voice graded A, or the Audiobook narrator blend, an even split with Delaney that we kept after listening. Play the calm narration sample on each card; a voice that sounds even over a paragraph will hold over a book.

  • How do I add pauses between paragraphs?

    Blank lines already produce a pause. For a longer one, insert a pause tag with the seconds you want, from 0.05 to 5. Tidy up script can also place pauses at long, unpunctuated stretches for you.

  • Can I slow down or emphasize part of the narration?

    Yes. The speed slider sets the overall pace from 0.5 to 2 times. A slow tag reads a single phrase at a reduced rate, which is the right tool for a line that needs weight.

  • Do the subtitles stay in sync on a long narration?

    Yes. The SRT and VTT files are built from the same render as the audio, with word-level timing on standard English voices, so the last cue is as accurate as the first.

  • Can I publish the narration in an audiobook or a video?

    Yes, on every plan including the free one. Standard voice output is licensed for commercial use, with no attribution line and no watermark.

  • How are the files delivered, and how long are they kept?

    Finished jobs land in My Speech with the MP3, SRT, and VTT. Paid plans can also download WAV. Free renders stay for 72 hours, paid renders for 30 days, and your text is deleted along with them.

Narrate your script in one take

Paste the whole piece, pick a narrative voice, and download the audio with subtitles.