RVC (Retrieval-based Voice Conversion)

Package

Self-hosted voice changer you train on 10 minutes of a voice, for speech or song

Price
Free, open source
Access
None, runs locally

About

MIT-licensed WebUI and CLI from the RVC Project. Train a model of one voice from about 10 minutes of clean audio, then convert recordings, singing or a live microphone into it (about 170 ms end to end). One trained model per voice; runs on NVIDIA, AMD or Intel GPUs or CPU. No pip package.

What you can do with it

  • Train a model of one voice and convert recorded speech into that voice
  • Make AI song covers by converting a sung vocal into another singer's voice
  • Run a live voice changer on a microphone with the real-time GUI or Windows VST plugin

Get started

  1. Install Python 3.12 and the requirements file that matches your GPU
  2. Download the HuBERT and RMVPE base models from Hugging Face into assets/
  3. Train a voice in the WebUI, then convert audio in the WebUI or infer/cli.py

Example

# Start the WebUI on port 7865 to train a voice model
python webui.py --noautoopen

# Convert a recording with a trained model
python infer/cli.py --model assets/weights/MODEL.pth \
  --input input.wav --output output.wav \
  --pitch 0 --f0-method rmvpe --index-rate 0.75

Details

Hosting
Runs locally, Self-hosted
Available in
Worldwide
MCP server
None

Tasks

Alternatives

Other tools for the same tasks.

Last checked on .