NNiceVois
One-tap RVC model training

Audio to PTH: train an RVC voice model from audio

Upload audio and start trainingUpload, choose epochs, name the model, and press Start Training once.
Summary

Upload your audio, choose the epochs, name your model, and press Start Training. NiceVois handles the complete RVC workflow and returns a portable .pth, matching .index, and ZIP.

Already have a .pth? This page is about making one. To convert a vocal with a model you already own, use the RVC voice converter.

Voice training

Train an RVC voice from an audio file

Add up to 50 WAV, MP3, FLAC, M4A, or OGG files. NiceVois keeps them separate and prepares them together as one private training dataset.

.pth

This name is used for the .pth, .index, and complete ZIP package.

Standard is selected automatically.
200 epochs

Recommended: add your audio and NiceVois will suggest the right number.

Clean audio produces a better model. Use a dry voice recording with minimal music, echo, or background noise. Prepare your audio first if it needs it.

Persistent cloud storage

Your model library

Your voice models are files you own. Download them once and no platform can take them away.

Loading…
Paid model packages stay saved until you delete them.

Your .pth, .index, ZIP, and private listening preview remain available after closing the browser or restarting your computer. A model from the free training run is kept for 7 days and shows its expiry date; download it before then. Full-resolution source audio and temporary training files are removed after validation.

Loading your models

Checking private cloud storage…

NiceVois is a converter-style utility for people who want the finished RVC files without installing a local training stack or maintaining a notebook.

What you can upload

The current public workspace accepts WAV, MP3, FLAC, M4A, and OGG. Use one speaker, minimal room echo, little or no music, and no clipping. A clean recording matters more than the filename extension.

Training from a song rather than a dry recording? Take the music off first, and take the harmonies off with it — a model trained on stacked background vocals learns the stack, not the singer. The free lead and backing vocal splitter returns the lead on its own.

What happens after upload

  1. The server receives the source audio and creates a private job for your browser.
  2. The RVC pipeline slices and prepares the dataset, extracts pitch and voice features, and waits for cloud GPU capacity.
  3. Training reports real completed epochs, then the service exports and validates the model.
  4. The matching retrieval index is built and both artifacts are packaged into a ZIP.

After the server accepts the job, training continues even if you close the tab or shut down your computer.

Which download should you keep?

FilePurposeBest use
.pthTrained RVC model weightsLoad the voice model into a compatible RVC interface.
.indexRetrieval index made from the training featuresUse with the matching model when your interface supports feature retrieval.
Complete ZIPThe matching .pth and .index togetherArchive or transfer the complete model safely.

Current public-beta limits

Your first training is free. Every job shows a fixed quote from audio length and epochs before cloud work begins, and may select up to 500 epochs. Purchased balance does not expire, finished downloads remain available for seven days, and new jobs may pause when cloud GPU capacity is full.

Turn your audio into a trained RVC model

Upload a supported voice recording, name the output, and follow the real training stages.

Open the audio-to-PTH converter

Audio to PTH questions

How do I convert audio to a .pth file?

Upload a voice recording (WAV, MP3, FLAC, M4A, or OGG), name the model, choose the epochs, and press Start Training. NiceVois runs the complete RVC v2 training in the cloud and returns a downloadable .pth, its matching .index, and a ZIP with both. No GPU, Colab notebook, or installation is involved.

Is converting audio to PTH free?

Your first run is free with a free account. It covers up to 150 epochs and under 5 minutes of audio, and you download the finished model before paying anything. Longer audio and stronger training up to 500 epochs are included in the monthly subscription, with a fixed quote shown before the job starts.

How much audio does a good RVC model need?

Two to ten minutes of clean, single-speaker audio is the practical range. Recordings under one minute may not produce a reliable model, and background music or overlapping voices hurt quality more than shortness does.

How long does audio to PTH training take?

Most jobs finish in 10 to 40 minutes depending on audio length and epochs. A live estimate is shown while the job runs, and you can opt into an email when the model is ready.

Related answers