Skip to content

ElevenLabs Music detection

The vocal is the interesting surface here, and it is also the hardest one to measure from a mixed file.

Output
Vocal-forward music generation
Lineage
Neural speech synthesis and voice cloning
Detection angle
High-band behaviour, indirectly
Attribution supported
No

Upload Audio or Drag & Drop

Upload a song to check whether its vocals or instrumental content show signs of AI generation. Analysis is performed by our AI music detection system.

  • MP3
  • WAV
  • FLAC
  • AAC
  • M4A
  • MP4
  • OGG
  • OPUS

MP3, WAV, FLAC, AAC, M4A, MP4, OGG, OPUS · max 25 MB (our upload limit) · 30+ seconds recommended

Your audio is uploaded over an encrypted connection and sent to our AI music detection partner solely to perform this analysis. We do not create a public report page and we do not intentionally retain your uploaded audio after the analysis completes. Upload only audio you are authorised to process — analysis does not transfer ownership or publishing rights.

  • Free
  • Fast
  • Secure
  • No registration

Vocoded audio in a finished mix

Neural vocoding leaves characteristic behaviour in the upper spectrum — sibilance that is too even, breath noise that repeats a little too exactly, transitions between phonemes that are cleaner than a human larynx manages. In an isolated vocal stem these are often audible to a trained ear.

In a mixed track they largely are not. Once drums, guitars and reverb occupy the same bands, a spectral measurement of the whole file sees the sum, not the voice. That is a genuine ceiling on what this engine can do, and it is why vocal-specific claims are kept out of the report.

It is worth separating three cases that get lumped together, because they raise different questions. A fully generated voice comes from a model with no specific person behind it. A converted voice is a real performance transformed towards a different timbre. A cloned voice is modelled on an identifiable artist. All three can appear in a track from a speech-synthesis lineage, and none of them is distinguishable from the others by acoustic measurement of a mix.

What the engine measures instead

The analysis works on the mix. It reads high-band energy distribution, spectral flatness, brightness variability over time and stereo behaviour, and reports how unusual the combination is. When a vocal is very forward and the backing is sparse, those readings carry more of the vocal's character and the result is more informative.

  • Sparse, vocal-forward material: the reading is more relevant to the voice.
  • Dense, loud production: the reading is dominated by the mix, not the vocal.
  • Any bitrate below roughly 192 kbps: the relevant band is largely gone.

What to listen for in the voice

Vocal cues are the one place where careful listening still outperforms measurement on a mixed file, because your ear can attend to the voice while a spectral average cannot. Listen on headphones at a moderate level and follow the vocal alone through a full verse and chorus.

Treat these as prompts for a documentation question, never as conclusions. Heavily processed human vocals produce several of them routinely.

  • Breaths that are missing entirely, or that repeat with an identical shape and length between phrases.
  • Sibilance that stays at a constant level regardless of how hard the line is sung.
  • Consonants that are cleaner than the room around them, as if edited in from a different take.
  • Emotional intensity that does not track the arrangement — the same delivery in a quiet verse and a big chorus.
  • Word stress that lands on the wrong syllable, or pronunciation that changes between repeats of the same line.

Isolate the vocal before you judge it

If the vocal is the question, get the vocal on its own. A dry, isolated stem removes the masking that defeats both your ears and the measurement, and it is the single most effective thing you can do to improve a reading on vocal-forward material.

There are two routes. Ask for it: anyone who actually recorded the part can supply a dry take in minutes, and an inability to produce one is itself informative. Failing that, a stem separation tool gets you close, with the caveat that separation introduces its own artefacts in exactly the high bands that matter — so a separated stem should raise your suspicion threshold, not lower it.

Analyse the isolated file as its own upload rather than reasoning about the mix. Sparse material gives the engine more of the voice, and the confidence label will reflect that.

  • Prefer a dry, unprocessed stem supplied by whoever recorded it.
  • Separated stems carry separation artefacts — read them more conservatively.
  • Avoid re-encoded copies; the informative band sits above most compression ceilings.
  • Run at least 30 seconds containing a full sung phrase, not a single held note.

Honest limitations

Heavy tuning, formant shifting and modern vocal chains make human voices measurably less human in exactly the bands that matter. A pop vocal through aggressive pitch correction and a de-esser can read like synthesis. We treat vocal-derived evidence as weak by design rather than letting it dominate the score.

The engine is also tuned for full music rather than for speech. A voice note, a podcast clip or a spoken deepfake is outside what these measurements were built to read, and the honest response to one of those files is that this is the wrong tool for the question.

And nothing here establishes whose voice it is. If the concern is that a recording imitates an identifiable performer, that is a consent and documentation question — ask who the target voice belongs to and for written permission — not something any acoustic probability can answer.

Getting a usable reading

Work from the highest-quality file available, prefer an isolated or sparsely accompanied vocal, and give the engine a full sung phrase rather than a fragment. Then read the confidence label first: on mixed, densely produced material the honest output is frequently inconclusive, and that is the result doing its job rather than failing at it.

ElevenLabs Music detection FAQ

  • Not reliably from a mixed file. Once the vocal sits inside a full arrangement, the measurements describe the mix rather than the voice. If you have an isolated vocal stem, the reading is considerably more informative.

Other generators