€5
Script to Voice Generator - Cartesia Sonic-3
Turn formatted scripts into fully voiced audio — powered by Cartesia Sonic-3, one of the most expressive life-like TTS engines on the market.
## Description
Script to Voice Generator converts formatted `.txt` or `.md` script files into fully voiced audio using **Cartesia Sonic-3** — a highly expressive cloud text-to-speech model with native emotion guidance and AI-generated laughter. Choose from **668+ voices** across multiple languages, accents, and genders.
Each speaker in your script gets their own voice, emotion, speed, and audio effects. Assign Radio, Reverb, Distortion, Telephone, Robot Voice, Cheap Mic, Underwater, Megaphone, Worn Tape, Intercom, Alien Voice, Cave, or Pitch Shift effects per character at Off / Mild / Medium / Strong levels. Two bonus toggles — FMSU (brutal digital corruption) and Reverse — round out the toolkit. Combine effects freely for distinct character identities.
Sonic-3 supports an **emotion system** with 56 values — set a baseline emotion per speaker in Tab 2, or override it per line with `{emotion: X}` directly in your script. The model performs the text with that emotional register rather than just reading it neutrally. Cartesia's Emotive voices produce outstanding results with the emotion system.
**`[laughter]`** placed anywhere in a dialogue line triggers AI-generated laughter — natural, voice-matched, inline. One of Sonic-3's most impressive features.
**Sonic-3 inline SSML tags** let you place mid-sentence pauses (`<break time="0.8s"/>`), inline speed/volume adjustments, and letter-by-letter spelling (`<spell>NATO</spell>`) directly in your script text.
**System requirements:** Windows 11 (tested). Windows 10 untested, try at your own risk. No Linux or macOS build available. Requires an internet connection and a Cartesia API key. FFMPEG must be on your system PATH.
**The generator produces:**
- Individual audio per each spoken line — clean (TTS only) and effects-processed versions
- All (effects-enabled) clips merged into a fully edited and smartly paced audio file, not normalized (true audio, better for media/games)
- All (effects-enabled) clips merged into a fully edited and smartly paced audio file, loudness normalized (even audio, better for podcasts)
- Reference .txt file with filenames, line numbers, and spoken content for every clip
**Features:**
- Cartesia Sonic-3 TTS — 668+ voices, expressive emotion guidance
- 56-value emotion system — per-speaker baseline, per-line override with `{emotion: X}`
- `[laughter]` nonverbalism — AI-generated inline laughter
- Sonic-3 inline SSML: `<break>`, `<speed>`, `<volume>`, `<spell>`
- Prosodic continuity — per-speaker context ID keeps voice consistent across lines
- Per-character voice, emotion, speed, pitch, volume, and audio effects
- 13 audio effects — most with Off / Mild / Medium / Strong presets
- 2 bonus toggles: FMSU (brutal corruption) and Reverse
- Yell Impact mode for punchy single-word exclamations
- Inner thoughts filter (Whisper, Dreamlike, Dissociated presets)
- Sound effect events (play/stop/loop) placed in the merge timeline
- Smart merged audio with configurable punctuation-based pause timing
- Loudness-normalized and raw merge outputs
- Per-line clip files (clean and effects) for editing flexibility
- Optional Pronunciation Dictionary ID support (fix mispronounced names/callsigns)
- Character profiles saved automatically between sessions
- Parse log with line-by-line error reporting
- Included example scripts and AI prompt templates to get started fast
**Ideal for:**
- Game dialogue
- Character oneliners
- Audio dramas
- Audio books
- Ambient narration
- Meditation audios
- Voice banks
...and any project where you want to convert a written script into audio.
---
### A note on Cartesia pricing
Cartesia is a **cloud API service** — you pay per model credit used. The free tier gives **20,000 credits per month**, which is a reasonable amount to get started with. Paid plans increase the limit significantly.
The quality-to-credit ratio is strong — Sonic-3 produces expressive, natural-sounding output that goes well beyond what typical TTS engines deliver. The recommendation is to use it **intentionally** — write tight scripts, test voices with the Test Voice button before full generation, and use the emotion system to get real performances out of the voices.
Check Cartesia pricing and manage your account at: https://play.cartesia.ai
---
*Includes in-depth guides covering script writing for Sonic-3 TTS, the emotion system, inline SSML tags, and the full audio effects pipeline.*
---
Check out everything else I do: ✨🚀
https://linktr.ee/reactorcore
€5
€5