Skip to content

Vocal Remover & Karaoke Maker

Split a stereo song into an instrumental backing and an isolated vocal, preview both, and download either as a WAV. The separation runs on your own device.

Drop a song to separate

or choose an audio file

MP3 · WAV · FLAC · AAC · M4A · MP4 · OGG · OPUS · up to 25 MB

Your audio stays on your device

This tool decodes, analyses and exports your track entirely in your browser. The file is never uploaded to our servers, and nothing is stored when you close the tab.

  • Free
  • No account
  • Browser processing

How the vocal remover works

Stereo mixing is a set of conventions as much as a technique, and one of the most consistent is that the lead vocal sits in the centre of the image. It gets there by being sent to the left and the right channel at equal level, which means that when you compare the two channels the vocal is the part that matches almost perfectly. Guitars, keys, room mics, reverb returns and stereo synths are deliberately placed away from that centre so the mix has width.

This tool exploits that. It slices the track into overlapping 4,096-sample frames, transforms each one into the frequency domain, and for every frequency bin compares the left and right channel using a normalised cross-power measure. A value near one means the two channels are carrying the same thing at the same level: centred content. A value near zero means they disagree: something panned, something wide, something stereo. That per-bin value becomes a soft mask, the mask is applied to build a centre signal, and the centre signal is subtracted from the original to leave the instrumental.

Using a soft mask rather than the old left-minus-right trick matters. Subtracting one channel from the other does remove the vocal, but it also collapses the mix to mono and cancels anything else that was centred, which is why those karaoke tracks always sound thin and lifeless. Masking in the frequency domain keeps the result in stereo and only touches the frequencies where centred energy actually exists.

Getting the best result from the controls

Separation strength is the main dial. At lower settings only the most unambiguously centred content is removed, which protects the drums and bass but leaves more vocal bleed. At higher settings the mask widens and catches more of the voice, at the cost of thinning the kick, snare and bass — the other three things that live in the middle of nearly every mix.

The frequency limits exist for the same reason. Below roughly 120 Hz there is almost never any vocal information, only kick and bass, so leaving that region untouched preserves the low end entirely. Above roughly 12 kHz you are mostly in cymbal and air territory, and processing it tends to add a swirling artefact without removing any perceptible vocal.

  • Start at the default strength and only raise it if you can still hear the lead clearly
  • If the drums lose punch, lower the strength or raise the low cut
  • Modern loud masters separate worse than dynamic mixes — the limiter glues everything together
  • Live recordings and mono-era records will not separate usefully at all

What people use this for

The obvious use is karaoke: pull the vocal out of a song you want to sing over. The second most common is the reverse — keeping the vocal to build an acapella for a remix, a mashup or a practice reference. Producers use it to check how much of a mix is centred, singers use it to study phrasing without the backing, and teachers use it to isolate a part for a lesson.

Whatever the use, remember that separating a commercial recording does not give you any right to distribute the result. The output is a derivative of a copyrighted master. Personal practice and private study are one thing; publishing an acapella or selling a karaoke track is another, and needs permission from the rights holders.

Frequently asked questions

  • Lead vocals are almost always mixed dead centre, meaning they appear at the same level and phase in both the left and right channel. The tool runs a short-time Fourier transform over the track and, for every frequency bin in every frame, measures how coherent the two channels are. Bins that behave like centred content are pulled into the vocal stem; everything else stays in the instrumental.