Text to Speech

Free online text-to-speech tool: natural main engine (CosyVoice 3) + emotional voice engine (IndexTTS-2), supports multiple voices, adjustable speed/pitch/volume, multilingual, emotional styles, voice cloning, and multi-speaker dialogue. Export MP3/WAV and SRT subtitles. GPU self-deployed for privacy protection.

Synthesis engine
0 / 2000
Loading preset voices…

Style/emotion works best with “Voice clone” and custom reference voices; preset named voices mostly keep their own character.

GPU synthesis is usually fast; the model loads first on first use or after idle (a few seconds).

FAQ

What model is used? Is it free?

Powered by self-deployed open-source dual engines CosyVoice 3 (natural main) and IndexTTS-2 (emotional voice), completely free. Offers daily free synthesis quota with no login required.

How to choose between the two engines?

"Natural Main" (CosyVoice 3) offers more natural, fluent speech and faster speed, with multiple preset voices—ideal for daily reading, customer service, and announcements; "Emotional Voice" (IndexTTS-2) delivers stronger, more dramatic expression—perfect for audiobooks, character voice acting, and video narration. Switch between them with one click at the top of the page.

Can I clone my own voice?

Yes. Select "Voice Cloning", upload a clear 5–15 second audio sample, enter corresponding text, and generate speech in your voice. Both engines support voice cloning. Reference audio is used only for this session and deleted immediately after, never stored.

Which languages and emotions are supported?

Supports Chinese, English, Japanese, Korean, Cantonese, and cross-language synthesis. The main engine allows natural language instructions in the "Style" field, such as "speak happily" or "in Sichuan dialect"; switching to the "Emotional Voice" mode lets you input emotions like "happy", "sad", "angry", or "surprised" directly for stronger emotional expression.

Why does synthesis take time?

GPU deployment ensures fast processing; first use or idle periods require initial model loading (a few seconds for main engine, 20–40 seconds for emotional engine). Long texts are synthesized sentence by sentence and streamed back—listen to completed parts while waiting.

Can I create multi-speaker dialogues/podcasts?

Yes. Enable "Multi-Speaker", input each character followed by their lines in separate lines, then assign different voices to each speaker.

Explore more AI tools and products