Skip to content

Accuracy and limitations

This page is the single place where we say what is known about how well this detector performs, how to read a result, and every way it is known to be wrong. There is no accuracy percentage on it, because we have not run the evaluation that would justify one.

Last updated August 2026 · engine build ensemble-1

Short answer

How accurate is AI music detection?

No one can honestly give a single figure, and we do not publish one. Accuracy varies with the generator, its version, the encode, the genre and how much human production followed. We publish the evaluation protocol with per-generator and per-encode arms instead, and every accuracy cell stays empty until it is measured rather than estimated.

Read the benchmark protocol →

Why there is no percentage here

Classification on this site is performed by a specialist third-party AI music detection service. We do not train or own a detection model, and we have not run a controlled evaluation of the current system against a labelled corpus. Any accuracy figure we quoted today would be marketing, not measurement.

This is why the report always pairs a probability with a confidence level, and why Inconclusive is returned readily. Those are not evasions. They are the correct behaviour for a system whose error rates we have not measured.

What would have to be true before we published one

  • A fixed evaluation dataset with documented sampling criteria
  • Separate development and test sets, with no leakage between them
  • Generator-held-out testing — evaluating on generators absent from tuning
  • Genre and language diversity, including instrumental-only material
  • Audio-quality variation: lossless, 320 kbps, 128 kbps, 64 kbps, resampled
  • A published confusion matrix with precision, recall and both error rates
  • A reproducibility note so others can check the claim

The full design is public at the benchmark protocol, and the state of every planned and completed study is tracked on the research hub.

What has and has not been measured

  • Measured: run-to-run determinism, how far an MP3 encode moves the score, how far a mono downmix or a ten-second excerpt moves it, and how often synthesiser-only human material reads as AI. These are published in the compression study and the electronic-music study.
  • Not measured: precision, recall, specificity, the false-positive rate on verified human recordings, the detection rate per generator, and behaviour on a held-out generator. These read not measured on the benchmark and are not estimated anywhere on this site.

How to read your result

The headline number is an estimate of how strongly your audio resembles patterns common in AI-generated music. It is not a confidence score and it is not a verdict. It is deliberately pulled toward the middle: a detector that hands out 99% is telling you about its own presentation choices, not about your song. Confidence is reported separately, and a high probability at low confidence is weaker evidence than a moderate probability at moderate confidence. The four labels, and what each one should and should not make you do, are explained on what your result means.

What the tool cannot do at all

  • It cannot prove how a track was made. Acoustic analysis observes properties of a signal. Authorship is a fact about a process, and no measurement of the output recovers it.
  • It cannot attribute a generator. There is no Suno, Udio or ElevenLabs classifier here. When evidence leans generated, the report says “unknown AI generator” and stops.
  • It cannot separate stems. A synthetic vocal over human instrumentation is measured as one mixed signal, not as two components.
  • It cannot look a song up. Nothing is matched against a database of known AI releases, so a widely circulated generated track gets no special treatment.
  • It cannot produce forensic evidence. The output is an estimate and is not admissible as proof in a copyright, disciplinary or employment process. See the disclaimer.

Known failure conditions

Lossy compression

Low-bitrate MP3 and AAC strip exactly the high-frequency detail that synthetic-audio detection leans on most heavily. A 128 kbps encode of a live band and a generated track can converge to a very similar signature. Our compression study measures how far the score moves across the encode ladder.

Heavy mastering and loudness maximisation

Modern mastering routinely pushes crest factor below 8 dB, removing the micro-dynamics that separate performed from generated audio. Generators were trained on loudness-maximised masters, so heavily mastered electronic and pop music is a known false-positive risk.

Human electronic music

Quantised, synthesiser-only, in-the-box human production reproduces most of the same statistics as generated audio. This is the single most common false-positive scenario, it is a genre bias rather than a threshold bug, and we measured it on our own reference renders: three of twelve read above 70%.

Post-produced generated tracks

The reverse failure. Re-recording, re-mixing, adding a live instrument, running the file through analogue gear, or simply re-encoding it several times removes most of what remains to be measured. Generated material that has been through a real production chain routinely reads as human.

Short clips

Below about 30 seconds there are too few independent windows to check for agreement. Five seconds is the hard minimum accepted; 45 seconds or more is meaningfully better. A ten-second excerpt moved the score by 4.4 percentage points on average in our own testing, and by up to 31 points in the worst case.

Unseen generators and generator drift

Published research on synthetic-audio detection consistently finds that strong test-set performance does not transfer to generators the detector was not built around. Each generator release also removes artefacts the previous one left behind, so any statement about detection is a statement about specific versions of both the generator and the engine.

Hybrid tracks

A human vocal over a generated backing, or a generated stem inside a human arrangement, produces exactly the window-level disagreement that triggers an inconclusive result. That is appropriate: the honest answer for a hybrid track is “partly”, which a single probability cannot express.

Material the tool was not designed for

Speech, applause, crowd noise, field recordings and near-silence are outside what the measurements were built for. For a spoken voice, read how to tell if a voice is AI generated instead.

False positives and false negatives

A false positive can damage a real musician’s reputation; a false negative merely misses a synthetic track. Those costs are not symmetrical, which is why the thresholds are conservative, why Likely AI-generated requires both a high score and agreeing segments, and why the false-positive rate on verified human recordings is the metric we consider most important in the benchmark.

Why independent validation matters

This page, the explainer and the benchmark protocol were written by the people who operate the detector. Self-evaluation has an obvious incentive problem, and no amount of careful wording removes it. The machine-readable facts are published at /benchmark.json under CC BY 4.0 precisely so that someone else can check them, and independent criticism will be published alongside our response.

How to get the most reliable reading

  • Use the highest-quality copy of the file you can obtain, ideally lossless.
  • Analyse 45 seconds or more of continuous music, and run two different excerpts.
  • Treat an inconclusive result as information, not as a failed attempt.
  • Weigh provenance above acoustics every time — the full process is on is this song AI generated.

How results should be used

  • As one input among several, never as the deciding factor
  • Alongside provenance: session files, stems, release history, artist conversation
  • Never as evidence in a copyright, employment, academic or enforcement decision
  • Never as the basis for a public accusation

See the detection-results disclaimer for the formal statement, read how the detection works, or go back and analyse a track.

Cite this page

Quotation with attribution is welcome. Please keep the wording of factual claims intact and link back to the source page.

AI Music Detector. “Accuracy and limitations of AI music detection.” Updated August 2026. https://aimusicdetector.co/accuracy