Dataset quality
Audio requirements for RVC voice training
The source recording teaches the model what to reproduce. Cleaner audio usually improves the result more than simply adding epochs to a weak dataset.
Supported upload formats
NiceVois currently accepts WAV, MP3, FLAC, M4A, and OGG. The service prepares the audio for the RVC training pipeline after upload, so you do not need to split the file manually.
Useful duration
The public beta accepts recordings up to five minutes. Use as much clean material as the limit permits. Under one minute can be useful for a fast experiment, but it may not contain enough variation for a dependable model.
What good source audio sounds like
- One clear target speaker at a time.
- Little or no music, room echo, crowd noise, or competing speech.
- No clipping, crackling, heavy distortion, or extreme compression.
- A reasonably consistent recording level.
- Natural variation in words, pitch, pace, and expression.
Problems the model can learn accidentally
RVC learns patterns from the supplied audio, including unwanted ones. Constant reverb can become part of the voice character. Instrumental bleed, harmonies, or another speaker can weaken identity. Aggressive noise removal can create metallic artifacts that the model then repeats.
| Source | Use it? | Why |
|---|---|---|
| Dry microphone recording | Best choice | Clear voice identity with little interference. |
| Clean isolated vocal stem | Usually good | Useful when separation did not leave music or watery artifacts. |
| Podcast or interview | Check carefully | Room tone, compression, or other speakers may be present. |
| Full mixed song | Avoid | Music and backing vocals become part of the training signal. |
| Several speakers in one file | Avoid | The model cannot reliably learn one consistent identity. |
Before uploading: listen through the recording with headphones. If you can clearly hear something you do not want in the finished voice, remove it or choose a cleaner source.
Upload it and let the interface calculate a duration-based training recommendation.
Open the audio uploader