RVC (Retrieval-based Voice Conversion)
PackageSelf-hosted voice changer you train on 10 minutes of a voice, for speech or song
- Price
- Free, open source
- Access
- None, runs locally
About
MIT-licensed WebUI and CLI from the RVC Project. Train a model of one voice from about 10 minutes of clean audio, then convert recordings, singing or a live microphone into it (about 170 ms end to end). One trained model per voice; runs on NVIDIA, AMD or Intel GPUs or CPU. No pip package.
What you can do with it
- Train a model of one voice and convert recorded speech into that voice
- Make AI song covers by converting a sung vocal into another singer's voice
- Run a live voice changer on a microphone with the real-time GUI or Windows VST plugin
Get started
- Install Python 3.12 and the requirements file that matches your GPU
- Download the HuBERT and RMVPE base models from Hugging Face into assets/
- Train a voice in the WebUI, then convert audio in the WebUI or infer/cli.py
Example
# Start the WebUI on port 7865 to train a voice model
python webui.py --noautoopen
# Convert a recording with a trained model
python infer/cli.py --model assets/weights/MODEL.pth \
--input input.wav --output output.wav \
--pitch 0 --f0-method rmvpe --index-rate 0.75Details
- Hosting
- Runs locally, Self-hosted
- Available in
- Worldwide
- MCP server
- None