AI vocal detection and its limitations
Our detector does not claim to detect AI vocals. This page explains what that would actually require — and why any tool promising it inside a finished mix deserves scepticism.
Last updated August 2026
Three different things people mean by “AI vocals”
- Fully generated vocals — the voice, melody and lyrics all come out of a music generation model. Nothing was sung.
- Voice cloning — a model trained on a specific person’s voice sings material that person never performed. This is the case with the most serious legal and ethical weight.
- Voice conversion — a real human sings the take, and a model transforms the timbre onto another voice. The performance, phrasing and breathing are genuinely human.
These leave very different traces. Conversion in particular preserves nearly all of the human performance cues a detector would look for, which is why it is the hardest of the three and the most commonly missed.
Why the mix makes it harder
In a finished master the vocal is compressed, de-essed, pitch-corrected, doubled, saturated, reverberated and sitting under two or three other layers. Every one of those steps removes or masks the fine spectral and temporal detail vocal detection depends on. By the time a track reaches a streaming platform, much of the evidence is gone.
What a credible vocal detector would need
- Source separation to isolate the vocal stem before analysis — and separation itself introduces artefacts that can be mistaken for synthesis
- A vocal-trained model, evaluated separately for generation, cloning and conversion
- Multilingual evaluation, since phonetics and singing style vary enormously
- Held-out voice testing, so performance is not measured on voices the model effectively memorised
- Robustness testing across mix processing and lossy encoding
The AI Music Detector engine does none of this. It applies no source separation, so it cannot make any vocal-specific claim, and it does not pretend to.
What our tool does instead
It measures whole-mix properties. A track with fully generated vocals may still shift those measurements, because generated vocals usually arrive with generated instrumentation — but that is an indirect inference about the track, not a finding about the voice.
If you suspect a cloned voice
Acoustic analysis should be your last resort, not your first. Voice cloning cases are usually resolved through provenance: who published it, when, from which account, with what distribution history, and whether the named artist confirms or denies it. Cloning a recognisable artist without permission may also raise personality, publicity and moral rights issues in many jurisdictions, independent of any detector output.
The general verification process is set out in how to check if a song is AI-generated, and the tool’s honest scope is on the accuracy page.