Voice conversion
Why AI vocals sound robotic — and what is actually missing
AI voice models follow the notes a person sings. But sounds like s, t, k and p have no note in them — they are little bursts of air and noise. So the model has almost nothing to follow, and gives back a soft, mushy version of them. The sung parts come out well. The bits in between do not. That gap is what people hear as “robotic” — and since the original recording still has those sounds, they can be put back.
It is usually put down to the model, or the audio, or not training for long enough. It is normally none of those. Once you know what to listen for, it is hard to unhear.
Listen to what survives and what does not
Play any AI cover and listen to the long held notes. They are usually good — the voice sounds like the right person, the tune is right, the wobble on the end of notes is there. Now listen to the edges of the words instead:
- The hissy sounds — s, sh, z, ch. They come back dull, or slightly lisped, or smudged into the word.
- The punchy sounds — t, k, p. They lose their snap, so words start soft instead of sharp.
- The breaths. The intake of air before a line often vanishes completely, and a lot of what makes a singer sound like a living person goes with it.
The sung notes are the easy part. Everything between them is where the illusion falls apart.
Why it happens
An AI voice model works by following two things: the note being sung, and what the target person's voice sounds like. It is genuinely good at that.
But an s is not a note. It is a hiss of air. A t is a tiny click. There is no tune in them to follow, so the part of the model that does the clever work has almost nothing to hold on to.
So those sounds do not really get converted. They get guessed at. And because they are quiet, nothing looks wrong — the file looks perfectly normal on screen. You only notice by listening.
Training for longer will not fix this. More training makes the voice sound more like the right person. It does not change the fact that there is no note to follow in an s. A model trained three times as long usually has the same soft s sounds, so retraining to chase it tends not to help.
The fix: put the real sounds back
Here is the thing that makes this solvable. You already have the sounds that went missing. They are sitting in the original recording you fed to the model. The singer said every s and took every breath properly. It is the conversion that lost them.
So rather than asking the model to do better, you take those bits from the original recording and drop them into the converted one. That is what a vocal engineer would do by hand, one word at a time. NiceVois does it automatically, and every conversion gives you two files: the plain result, and the repaired one with the s sounds and breaths back in.
Why it is harder than it sounds
Copying a piece of audio from one file into another is easy. Making it sound like it belongs there is not. Four things decide whether the repair works or ends up worse than what it replaced:
- Finding the right bits. You cannot search for them by note, because they have no note. They have to be found by their hissy, noisy character instead.
- Lining them up. The converted version drifts against the original by tiny amounts. If the timing is off even slightly, the repair lands on the wrong sound and makes things worse.
- Matching each piece to the exact sound it replaces. Not to the singing around it. A conversion damages an s far more than it damages a held note, so if you match against the singing you badly underestimate the gap you are trying to close.
- Cutting the end at the right moment. The most obvious failure is a thud at the end of a repaired piece, where the original voice bleeds back in. The cut has to happen right where the next sound starts.
All four came from sitting and listening to real vocals. Settings that looked right written down were often wrong in the ear.
Where it still struggles
This works well most of the time, and it is not perfect. The hard case is when the original singer and the converted voice are very different from each other — a big jump in range, or a much brighter or darker voice. The repair takes a piece from one and puts it in the other, so the further apart they are, the more of a stretch that is, and you may still hear the joins. It is the part we are still working on.
How to give it the best chance
The repair can only be as good as the recording it borrows from. Two things genuinely help:
- Convert a vocal on its own, never a full song. Split the lead vocal out first — and if the song has harmonies, separate those too, or they get converted as well and the voice ends up singing against itself.
- Do not over-clean the source. Heavy noise reduction, or a tool that tames s sounds, will strip out the very things the repair needs to borrow. A slightly rough recording is better than a scrubbed one.
What about echo? That one is a judgement call, and it has its own guide — including a trick advanced users apply when they want the character of the original and clean sounds to borrow from.
Common questions
Why do AI vocals sound robotic?
Mostly the missing s and t sounds and the missing breaths. Voice models handle sung notes well and noisy little sounds badly.
Would a better recording fix it?
Better audio helps the voice sound more like the right person. It does not change how the model handles sounds with no note in them.
Can I do the repair myself?
Yes. It is a hand technique from vocal production, and doing it across a whole song is slow but perfectly possible. NiceVois automates it.
Does this apply to speech too?
Yes. It is just more obvious in singing, where long held notes sit right next to sharp little sounds.
Every conversion gives you the plain result and the repaired one, so you can compare them on your own audio.
Try the voice converter