Short-form voiceover
Hooks and one-line reads over vertical video, the use our note names. Word-level SRT for word-by-word captions.

Short-form voiceover is what our note gives Giulia, one of two Italian voices on the published list. Her record is a C grade at B target quality from 10 to 100 minutes of training audio, out of an Italian total of one to ten hours, and she returns word-level timing for captions.
Hear it
On the record
Short-form voiceover means a line or two over a vertical video: a hook, a caption read aloud, a product tag. Word-level timing lets the subtitle pop word by word, which is the style those clips use.
The Italian record is the same for both voices, tens of minutes at B target quality, a C grade. It measures training data, not delivery. The second sample, the ten-second hook, is the closest test of what our note asks of her.
Length the engine likes. The published guidance puts the sweet spot at roughly 100 to 200 tokens, a few sentences; very short lines are the weak spot and very long ones can rush.
Where we would use it
Hooks and one-line reads over vertical video, the use our note names. Word-level SRT for word-by-word captions.
Short promotional lines where one sentence is the whole script.
Lorenzo is the other Italian voice; the pairings below start with him.
Blends
No recipe we kept includes Giulia. Each card is a pairing to try; she keeps the Italian pronunciation and the second voice lends timbre.
A pairing to try, two Italian voices: Lorenzo reads tutorial scripts by our note.
50% Lorenzo
A pairing to try. Sienna is bright by our note and our hook voice in English; listen for that push on Italian.
50% Sienna
A pairing to try. Delaney's throaty edge sits behind kept recipes in four languages; listen for it under Italian.
60% Delaney
A pairing to try, two Romance-language voices: Elodie's note is a measured delivery.
50% Elodie
A fixed sample at this mix, rendered once and cached. No credits used.
Every blend works the same way: the first voice sets the pronunciation and pacing, the second voice lends only its timbre, and you choose the ratio between 10 and 90 percent. See the three steps on the voice library page
At a glance
| Language | Italian |
|---|---|
| Voice | Female |
| Overall grade | C |
| Target quality | B |
| Training audio | 10 to 100 minutes |
| Captions | Word-level captions |
| Cost | 1 credit per job, room for 10,000 characters |
| Formats | MP3 free, WAV on paid plans |
| Commercial use | Included on every plan |
| Preview | Free, no sign-up |
Grade, target quality, and training audio come from the engine's published voice list and estimate training data, not how a voice sounds.
FAQ
That is the use our note names. The TikTok page sets up the short-form workflow and the word-timed captions that style needs.
Yes. Italian returns word-level timing, so the SRT and WebVTT can highlight word by word.
Billing counts UTF-8 bytes, and Italian is mostly one byte per character with accented letters at two, so a credit carries close to 10,000 characters in one job.
Ten to a hundred minutes of training audio at B target quality. The grade measures data, not how she sounds.
Paste a script, keep Giulia or blend a second voice, and download the MP3 with subtitles.