CosyVoice 2
CosyVoice 2 is Alibaba's second-generation voice synthesis model, combining multilingual text-to-speech with voice cloning capabilities. Developed at Alibaba DAMO Academy, it supports a wide range of languages and is designed to produce natural, expressive speech across diverse speaker profiles, including zero-shot voice cloning from short audio references.
The model improves on its predecessor with enhanced prosody, better cross-lingual speaker consistency, and refined cloning fidelity. CosyVoice 2 is suited for multilingual content production, personalized voice experiences, and developers building voice-enabled applications across Asian and global language markets.
Key Features
Multilingual TTS covering a broad range of languages including Chinese and English
Zero-shot voice cloning from short audio reference samples
Improved prosody and natural intonation over CosyVoice 1
Cross-lingual speaker consistency for multilingual content
Expressive speech with controllable style and emotion
Ideal Use Cases
Multilingual audio content production for global markets
Personalized voice cloning for branded audio experiences
Localization of audio content into multiple languages with consistent voices
Voice-enabled applications targeting Asian language users
Automated narration with a consistent cloned speaker voice
Example Prompts for CosyVoice 2
Technical Specifications
| Provider | CosyVoice |
| Category | Audio |
| Modality | Text -> Audio |
Frequently Asked Questions
Try CosyVoice 2 now
Start using CosyVoice 2 instantly — 100 free credits, no credit card required. Access 750+ AI models through one platform.