ComfyUI custom node integrating VoxCPM — a tokenizer-free TTS system for expressive speech generation and voice cloning.
VoxCPM models speech in a continuous space using a MiniCPM-4 backbone, producing highly expressive speech and accurate zero-shot voice cloning. This node handles model downloading, memory management, and audio processing end-to-end.
- Voice Design — generate voices from natural language descriptions
- Controllable Voice Cloning — clone with style control instructions
- Ultimate Cloning — combine reference audio (identity) + prompt audio (prosody)
- 48kHz output, 30+ languages
- 44.1kHz output, LoRA support, native LoRA training
- Context-aware expressive speech, zero-shot TTS
- Automatic model management
Via ComfyUI Manager: Search ComfyUI-VoxCPM → Install.
Manual install:
cd ComfyUI/custom_nodes/
git clone https://github.com/wildminder/ComfyUI-VoxCPM.git
cd ComfyUI-VoxCPM
pip install -r requirements.txtRestart ComfyUI. Nodes appear under audio/tts. Models auto-download to ComfyUI/models/tts/VoxCPM/ on first use.
| Model | Params | Sample Rate | Languages | Link |
|---|---|---|---|---|
| VoxCPM2 | 2B | 48kHz | 30+ | openbmb/VoxCPM2 |
| VoxCPM1.5 | 800M | 44.1kHz | 2 | openbmb/VoxCPM1.5 |
| VoxCPM-0.5B | 640M | 16kHz | 2 | openbmb/VoxCPM-0.5B |
Unified TTS node supporting VoxCPM1.5 and VoxCPM2: zero-shot TTS, voice design, voice cloning, ultimate cloning, LoRA support.
Configures audio-based cloning (prompt/reference audio, VAD trimming). Connect to TTS node's voice_config input.
Configures diffusion parameters (temperature, sway sampling, CFG, timesteps, retry). Connect to TTS node's advanced_params input.
VoxCPM Train Config— LoRA training parametersVoxCPM Dataset Maker— create training datasets from audioVoxCPM LoRA Trainer— train custom LoRA models
Note:
voice_designis a direct parameter on the TTS node, not part of the Voice Cloning config.
Config precedence: Direct parameters > config node values > defaults.
Add VoxCPM TTS → select model → enter text → generate.
Connect Load Audio → prompt_audio, provide exact transcript in prompt_text → generate.
Select VoxCPM2 model → enter description in voice_design (e.g., "warm female voice") → generate. Voice design is applied in plain TTS and reference cloning modes. Ignored when prompt audio is used (continuation cloning).
Connect reference audio to reference_audio (no transcript needed) → generate. Voice design instructions can be combined with reference audio for controllable cloning (e.g., style control).
Connect reference_audio (identity) + prompt_audio with transcript (prosody) → generate.
Note
Denoising: The built-in ZipEnhancer denoiser is disabled by default to keep dependencies light.
| Parameter | Default | Range | Description |
|---|---|---|---|
temperature |
1.0 | 0.1-2.0 | Lower = stable, higher = expressive |
sway_sampling_coef |
1.0 | 0.0-2.0 | Sway sampling trajectory |
use_cfg_zero_star |
True | — | CFG-Zero* optimization |
cfg_value |
2.0 | 0.1-10.0 | Guidance scale |
inference_timesteps |
10 | 1-100 | More steps = higher quality, slower |
| Parameter | Default | Options | Description |
|---|---|---|---|
device |
auto | cuda, cpu, mps, xpu, npu | Inference device |
dtype |
auto | auto, bf16, fp16, fp32 | Model precision |
AudioVAE always runs in FP32 for numerical stability.
| Parameter | Default | Range | Description |
|---|---|---|---|
trim_silence |
False | — | VAD silence trimming |
max_silence_ms |
200.0 | 0-1000 | Max silence at boundaries (ms) |
top_db |
35.0 | 10-60 | Lower = more aggressive trimming |
Inference: Place .safetensors LoRA files in ComfyUI/models/loras/, refresh, select in lora_name dropdown.
Training: 👉 Full LoRA Training Guide
| Description | Result |
|---|---|
warm female voice |
Soft, gentle female voice |
deep male voice |
Low-pitched male voice |
cheerful young girl |
Energetic, high-pitched |
professional announcer |
Clear, authoritative |
whispering voice |
Quiet, intimate |
Combine descriptions: "warm female voice with slight British accent"
- Verbatim transcript —
prompt_textmust match audio word-for-word - Punctuation matters — affects intonation
- 5-15 seconds of clear speech works best
Warning
prompt_text is the exact transcript, not a description of the voice.
- Voice cloning can be misused for deepfakes — use responsibly
- May exhibit instability with very long/complex inputs
- VoxCPM1.5: Chinese and English only; VoxCPM2: 30+ languages
- Custom model selector UI — replaced default LiteGraph dropdown with custom DOM widget featuring cyber-themed design, SVG model icons, and DEFAULT/CUSTOM badges
- Real-time download progress — live progress bar with cancel support, Xet dedup tracking, and file-level progress
- Lazy-loaded frontend — 95% startup payload reduction; heavy JS loads only when a VoxCPM node is placed
- Model directory dialog — browse and select custom model paths via a dedicated dialog instead of queue-time prompts
- Architecture version tags — model dropdown detects and displays architecture version (v1/v2) tags
- ComfyUI settings API migration — model path settings now use ComfyUI's native settings panel
- Graceful download error handling — automatic retry on transient network errors with toast notifications
- BEM-styled UI — modern CSS architecture with design tokens for consistent theming
- VoxCPM2 support (voice design, reference cloning, ultimate cloning, 48kHz, 30+ languages)
- Unified dtype/device handling delegating to ComfyUI
- Vite 5→8 upgrade, LiteGraph canvas widget fixes
- LoRA training nodes
VoxCPM model and components: Apache-2.0 License by OpenBMB.
- OpenBMB & ModelBest for VoxCPM
- The ComfyUI team for the platform
══════════════════════════════════
