NNiceVois

AI covers

Why you separate the lead vocal before making an AI cover

Summary

A voice model converts everything you give it. If the backing vocals are still in the file, they get converted too — so the target voice comes back singing its own harmony stack against itself. That is what people are hearing when an AI cover sounds doubled, phasey or strangely crowded in the chorus. The fix is to convert the lead alone and mix the original harmonies back underneath.

It is a common reason an otherwise correct AI cover still sounds wrong, and it has nothing to do with the model or how long it was trained.

What actually happens

Say you take a song, run it through a vocal remover, and feed the vocals file to your voice model. That file contains the main vocal, but it also contains the doubles on the chorus, the harmony a third above, and every ad-lib in the last minute.

The model has no idea any of that is there. It converts all of it. The harmony that used to be a different singer sitting behind the lead now sounds like exactly the same person, because both have been turned into the target voice. Two copies of one voice, a few notes apart, landing at slightly different moments — that is precisely the smeared, doubled sound people describe.

What we found when we tried it: doing it in one pass on a real cover put every backing vocal through the voice model alongside the main one, and the harmonies came back as a second copy of the target voice. That is what sent us to the two-pass approach.

What to do instead

  1. Split the song into vocals and instrumental. One pass, and every tool does this.
  2. Split the vocals again into lead and backing. A second pass, over just that vocals file. This is the step most workflows skip.
  3. Convert only the lead. One voice in, one voice out.
  4. Mix back the original backing vocals and the instrumental underneath the converted lead.

The harmonies stay in the original singer's voice, which is how a lot of covers are put together anyway.

When you genuinely do not need it

Not every song has harmonies. If the vocal is one line the whole way through — a lot of rap, plenty of singer-songwriter tracks, most demos and voice notes — the ordinary vocals file is fine, and the extra step only costs you time.

The tell is the chorus. If the chorus suddenly sounds fuller than the verse, there are almost certainly doubles or harmonies in it.

The other two things that ruin a source vocal

This is the first of three. The other two matter for the same reason — a voice model converts whatever is in the file, including things that are not the singer:

Common questions

Why does my AI cover sound doubled?

Almost always because the harmonies were converted along with the lead. Separate the lead first and the doubling disappears.

Can I just turn the backing vocals down?

Not once they have been converted — by then they are baked into the same file as the lead. Separate before converting, not after.

Does this apply to speech as well as singing?

The harmony problem is specific to music. But the general rule holds everywhere: convert only the voice you want, with as little else in the file as possible.

Split the lead out first

Lead, backing and instrumental from one upload. Ten free minutes to try it.

Open the stem separator