XTTS v2 is a cross-lingual text-to-speech model from Coqui AI's XTTS line, supporting voice cloning and synthesis across 16 languages from a short audio reference. It allows a speaker's voice characteristics to be transferred to text output in any supported language, including languages different from the reference clip's original language.
XTTS v2 is particularly valued in the open-source TTS community for its cross-lingual cloning capability, which enables consistent speaker identity across language boundaries — useful for dubbing, multilingual content creation, and personalized voice experiences. It runs locally on consumer hardware and is widely used in developer toolkits and TTS pipelines.
Key Features
Cross-lingual voice cloning across 16 supported languages
Voice cloning from short audio reference samples
Speaker identity preserved when synthesizing in a different language
Local deployment on consumer GPU hardware
Open-source with active community integrations and tooling
Natural speech quality with consistent prosody across languages
Ideal Use Cases
Multilingual dubbing with consistent speaker voice identity
Personalized cross-lingual narration for international content
Voice cloning for content creators working in multiple languages
Developer integrations requiring local, API-free TTS with cloning
Accessibility solutions delivering content in a user's preferred language with a familiar voice
Example Prompts for XTTS v2
Technical Specifications
| Provider | XTTS |
| Category | Audio |
| Modality | Text -> Audio |
Frequently Asked Questions
Try XTTS v2 now
Start using XTTS v2 instantly — 100 free credits, no credit card required. Access 750+ AI models through one platform.