
VoxCPM2
VoxCPM2 Multilingual Text-to-Speech Voice Design & Controllable Voice Cloning
Text to Speech
Choose a voice, write your script, and generate a clean audio take.
VoxCPM2 Core Capabilities
Explore the model capabilities behind multilingual TTS, voice design, and controllable voice cloning.
2B Multilingual Speech Model
VoxCPM2 supports text-to-speech in 30 languages and nine Chinese dialects, with a MiniCPM-4 backbone.
Tokenizer-Free Speech Generation
The model avoids external discrete speech tokenizers and generates continuous speech representations; text is still encoded by the language model tokenizer.
Natural-Language Voice Design
Describe a voice in natural language and generate a new voice without providing reference audio.
Controllable Voice Cloning
Provide reference audio and optional style guidance to preserve timbre while adjusting pace, emotion, or expression.
48 kHz Audio Output
VoxCPM2 accepts 16 kHz reference audio and produces 48 kHz output through its AudioVAE pipeline.
VoxCPM2 Model Facts
Official model facts are shown below; online latency and audio quality depend on deployment settings.
Latest VoxCPM2 release
Plus 9 Chinese dialects
Supports 16 kHz reference input
Official PyTorch reference
Source: OpenBMB VoxCPM2 model documentation; service results may differ.
VoxCPM Model Versions
Compare official facts for the current VoxCPM2 release and earlier VoxCPM checkpoints.
| Model | Audio output | Parameters | Reference audio | Zero-Shot | Multilingual | Open Source |
|---|---|---|---|---|---|---|
VoxCPM2 Current VoxCPM2 release used by this service | 48 kHz | 2B | 16 kHz input | |||
VoxCPM1.5 Earlier VoxCPM1.5 release | 44.1 kHz | 0.6B | Reference + continuation | |||
VoxCPM-0.5B Legacy VoxCPM checkpoint | 16 kHz | 0.5B | Continuation |
These are official model-version facts; service latency and audio quality depend on deployment.
VoxCPM2 Audio Demos
Explore cross-language speech, emotion, and context-aware generation with the online demo.
Cross-Language - EN→CN
English speaker voice cloned to Chinese speech
Cross-Language - CN→EN
Chinese speaker voice cloned to English speech
Emotion - Happy
Emotionally rich happy tone expression
Emotion - Sad
Emotionally rich sad tone expression
Context-Aware - News
Intelligent news broadcasting style
Context-Aware - Story
Intelligent storytelling style
More audio samples available at Official Demo Page
VoxCPM2 Technical Resources
Read the VoxCPM2 technical report and open-source implementation, then try the official model resources.
Academic Paper
Detailed technical principles, experimental results, and performance evaluation reports
Read PaperQuick Start
Install the official package and generate multilingual speech with the VoxCPM2 model.
Get StartedTechnical Highlights
Application Scenarios
VoxCPM2 supports practical workflows across content creation, learning, accessibility, and multilingual audio production.
Audiobook Production
Rapidly generate high-quality audiobooks with consistent voice
Language Learning
Personalized speech education with multilingual accent training
Content Creation
Professional voice solutions for video dubbing and podcast production
Accessibility Applications
Personalized reading experiences for visually impaired individuals
Run VoxCPM2 Locally
Install the official package and generate multilingual speech with the VoxCPM2 model.
Install the package
Install the official Python package (Python ≥ 3.10, PyTorch ≥ 2.5, CUDA ≥ 12.0).
Load the VoxCPM2 model
Load the latest weights from Hugging Face:
Generate audio
Generate speech and save the native sample rate:
System Requirements
Frequently Asked Questions
Common questions and answers about VoxCPM2 technology and usage