Why we are featuring it
Synthetic speech that does not start from discrete tokens.
VoxCPM generates continuous speech representations through a diffusion autoregressive architecture. VoxCPM2 combines languages, voice design, cloning, and 48kHz audio in a model usable through Python and compatible serving options.
What is inside
- The 2B-parameter VoxCPM2 model supports 30 languages.
- Voice design from a natural-language description.
- Controllable and ultimate cloning from reference audio.
- 48kHz output and real-time streaming with vLLM options.