Limitations
Short answer: this detector estimates how strongly a recording resembles AI-generated audio. It cannot prove authorship, it cannot name a generator, and it is wrong in both directions often enough that no single result should decide anything. This page lists every failure mode we know about, in one place.
Last updated August 2026 · engine build ensemble-1
What the tool cannot do at all
- It cannot prove how a track was made. Acoustic analysis observes properties of a signal. Authorship is a fact about a process, and no measurement of the output recovers it.
- It cannot attribute a generator. There is no Suno, Udio or ElevenLabs classifier here. When evidence leans generated, the report says “unknown AI generator” and stops.
- It cannot separate stems. A synthetic vocal over human instrumentation is measured as one mixed signal, not as two components.
- It cannot look a song up. Nothing is matched against a database of known AI releases, so a widely circulated generated track gets no special treatment.
- It cannot produce forensic evidence. The output is an estimate, and it is not admissible as proof in a copyright, disciplinary or employment process. See the disclaimer.
Where it goes wrong most often
Lossy compression
Several measurements are computed in exactly the frequency region an encoder discards first. A 96 kbps MP3 of a live jazz quartet can present the same hard spectral ceiling as generated audio. The tool lowers its own confidence when it detects that kind of degradation, but degradation it cannot detect still moves the number. Our compression study sets out the mechanism.
Heavy mastering and loudness maximisation
Limiting compresses crest factor and micro-dynamics — two of the properties that distinguish performed from generated audio. A modern commercial master is therefore harder to read than a rough mix, and errs toward the generated end.
Human electronic music
Quantised, synthesiser-only, in-the-box human production reproduces most of the same statistics as generated audio: uniform tonal balance, consistent stereo width, little articulation change. This is the single most common false-positive scenario, and it is a genre bias, not a bug we can fix with a threshold.
Post-produced generated tracks
The reverse failure. Re-recording, re-mixing, adding a live instrument, running the file through analogue gear, or simply re-encoding it several times removes most of what remains to be measured. Generated material that has been through a real production chain routinely reads as human.
Generator drift
Each generator release removes artefacts the previous one left behind. Any statement about detection is a statement about a specific generator version measured against a specific engine build; it says little about the next release of either.
Hybrid tracks
A human arrangement built around a generated stem belongs to neither class, and there is no correct binary answer. Sections of the same file may disagree with each other, which usually produces an inconclusive result — the honest outcome rather than a malfunction.
Short or unusual audio
Below roughly ten seconds there is not enough material to sample multiple sections and test whether they agree, so confidence collapses. Speech, applause, crowd noise, field recordings and near-silence are all outside what the measurements were designed for.
What the numbers do and do not mean
The reported probability is capped between 15% and 85%, shrunk toward the middle, and paired with a separate confidence level. That cap is deliberate: no file-only measurement justifies certainty. A high probability with low confidence is a weaker signal than a moderate probability with high confidence, and the two should never be collapsed into one figure.
We publish no accuracy percentage, because we have not completed the evaluation that would justify one. The protocol, corpus design and metric definitions are public at /benchmark, with every result cell currently null. Nulls mean unmeasured; they are never filled with estimates.
How to compensate
- Use the highest-quality copy of the file you can obtain, ideally lossless.
- Analyse 45 seconds or more of continuous music, and run two different excerpts.
- Treat an inconclusive result as information, not as a failed attempt.
- Weigh provenance above acoustics every time — the full process is on is this song AI generated.
- Read what your result means before acting on a number.
Conflict of interest
This page, the methodology and the benchmark protocol were all written by the people who built the detector. That is a real limitation of its own. Independent evaluation is welcome: the machine-readable facts are at /benchmark.json under CC BY 4.0, and contact is open.
Where to go next
- Detector changelog — which build produced your result.
- How to evaluate an accuracy claim — including ours.
Cite this page
Quotation with attribution is welcome. Please keep the wording of factual claims intact and link back to the source page.
AI Music Detector. “Limitations of AI music detection.” Updated August 2026. https://aimusicdetector.co/limitations