How an AI Cover Changes the Voice
After hearing an AI cover in Sun Yanzi’s voice, I wanted to try making “Zhi Wo” sound like JJ Lin. I used RVC, a tool for converting vocal timbre. I couldn’t find a suitable ready-made model, so I prepared to train one.
The song and samples were ready. Once training started, I found out just how slow it was. I never finished the cover.
The performance is already there
This kind of cover needs two different sets of vocals. JJ Lin’s recordings provide the target voice for training. The original vocals from “Zhi Wo” are the input for conversion. They already contain the words, melody, pauses, and held notes; the model doesn’t have to compose a song from lyrics.
If the original singer holds a syllable for two seconds, those two seconds of audio are what gets converted. The original timing and pronunciation still shape the result. Sounding like a particular singer doesn’t mean that singer would actually perform the song that way.
The vocals also need to be separated from the backing track. After conversion, the new vocals are mixed with the original accompaniment. Both training and conversion need vocals that are as clean as possible; the accompaniment is kept for the final mix.
The notes stay in place and change color to illustrate a change in vocal timbre.
Turning recordings into training data
My notes list four songs, about eighteen minutes in total. Song length and usable vocal length are different things. Intros and instrumental passages contain no singing, while separated vocals may still contain instruments or someone else’s harmonies. Eighteen minutes of songs isn’t eighteen minutes of clean audio from the target singer.
RVC’s official README (opens in a new tab) recommends low-noise recordings. Collecting enough minutes is only part of preparing the samples; separating vocals doesn’t automatically make them clean.
The training page then processes the audio. After selecting the input folder and output sample rate, RVC reads the recordings, splits them into short clips, and saves audio for training. It also saves a 16 kHz copy for feature extraction. These files come from the same recording but serve different purposes. Renaming the original file won’t do it. The preprocessing code (opens in a new tab) handles slicing, level adjustment, and resampling.
Next, the audio has to become numbers the model can use. HuBERT extracts audio features. With pitch guidance enabled, RVC also extracts F0, the fundamental frequency at each point in time. Think of it as the changing pitch: the same syllable can be sung high or low, and conversion needs to know which note was sung. Feature extraction (opens in a new tab) and pitch extraction (opens in a new tab) are separate steps.
The model and the retrieval index
Training needs the audio clips and their matching features. RVC pairs those files and starts training with the selected model version, sample rate, and pitch settings. Pretrained weights must match those choices. The epoch count sets how many training passes to run; batch size sets how many examples are processed together. The official WebUI (opens in a new tab) separates preprocessing, feature extraction, and model training.
An epoch is roughly one pass through the training data, during which the model is adjusted based on its errors. Training for 150 epochs means repeatedly using the same samples, rather than collecting 150 new sets of recordings. The count affects how long training takes, but doesn’t by itself tell you how convincing the voice will be. The training code (opens in a new tab) has separate loops for epochs and for reading the data within each epoch.
RVC also has a retrieval index. The trained model is a .pth file; the index is an .index file that organizes features from the training recordings for similarity searches during conversion. The model generates audio, while retrieval supplies reference features from the target voice. Building the index is a separate operation, rather than another pass of model training. The index code (opens in a new tab) uses the features extracted earlier.
Only at conversion time do the “Zhi Wo” vocals enter the process: their features and pitch are extracted, the target voice model converts them, and the result is mixed with the backing track. Using the model to produce audio is a different computation from training it on JJ Lin’s voice.
Everything was ready, but training would take hours
That’s where I stopped. My notes put CPU training at five to six minutes per epoch, with 150 epochs planned. I asked ChatGPT how long it would take. At that pace, the full run worked out to roughly twelve and a half to fifteen hours. That was an estimate at the time. I never completed the run, and it wasn’t the time RVC needed to convert a single song.
Time wasn’t the only problem. I didn’t know how convincing the trained voice would sound. The environment, samples, and intermediate files also took up a lot of disk space. Hours of waiting for an uncertain result took all the excitement out of it.
I gave up waiting. I never finished the model or produced the cover.
Comments