Kokoro 82M is an extremely compact text-to-speech model with only 82 million parameters, purpose-built for low-latency, on-device voice synthesis. Its minimal footprint makes it deployable on edge hardware, mobile devices, or embedded systems where larger TTS models would be impractical.
Despite its size, Kokoro 82M produces intelligible, natural-sounding speech across a range of voices. It is optimized primarily for speed and efficiency rather than studio-grade fidelity, making it well-suited for applications that need real-time speech output without a round-trip to a cloud API.
Key Features
Ultra-compact 82M parameter architecture for on-device deployment
Very low inference latency enabling near real-time voice output
Multiple voice style options within a single lightweight model
Suitable for CPU-only inference without GPU requirements
Designed for edge, mobile, and embedded runtime environments
Ideal Use Cases
Screen readers and accessibility tools running locally on-device
Voice-enabled mobile apps needing offline TTS capability
IoT and embedded device voice prompts
Rapid prototyping of voice interfaces without cloud dependency
Low-resource server environments requiring high-throughput TTS
Example Prompts for Kokoro 82M
Technical Specifications
| Provider | Kokoro |
| Category | Audio |
| Modality | Text -> Audio |
Frequently Asked Questions
Try Kokoro 82M now
Start using Kokoro 82M instantly — 100 free credits, no credit card required. Access 750+ AI models through one platform.