Skip to content

Accuracy, testing and limitations

We do not publish an accuracy percentage, because we have not yet run an evaluation that would justify one. This page explains that decision and what we are doing about it.

Last updated August 2026

The short version

The AI Music Detector analysis engine is transparent about what it measures. It measures real properties of your audio and reports how they compare with patterns often seen in generated recordings. It has not been validated against a labelled dataset with held-out generators, so any accuracy figure we quoted today would be marketing, not measurement.

This is why the tool caps its own output at 15–85%, shrinks scores toward the middle and returns Inconclusive readily. Those are not bugs. They are the correct behaviour for an uncalibrated detector.

What we would need before publishing a number

  • A fixed evaluation dataset with documented sampling criteria
  • Separate development and test sets, with no leakage between them
  • Generator-held-out testing — evaluating on generators absent from tuning
  • Genre and language diversity, including instrumental-only material
  • Audio-quality variation: lossless, 320 kbps, 128 kbps, 64 kbps, resampled
  • A published confusion matrix with precision, recall and both error rates
  • A reproducibility note so others can check the claim

Until all of that exists and is published on the research page, the correct answer to “how accurate is it?” is: unknown, and probably modest.

Known failure conditions

Lossy compression

Low-bitrate MP3 and AAC remove exactly the high-frequency detail two of our measurements rely on. A 128 kbps encode of a live band can look identical to a generated track on the spectral-ceiling measurement. The tool flags this by downgrading audio quality to Limited or Poor, which in turn caps confidence.

Loudness-maximised mastering

Modern mastering routinely pushes crest factor below 8 dB. That resembles generated output, so heavily mastered electronic and pop music is a known false-positive risk.

Short clips

Below about 30 seconds there are too few independent segments to check agreement, which is the engine’s strongest signal. Ten seconds is the hard minimum; 45 seconds or more is meaningfully better.

Unseen generators

Published research on synthetic-audio detection consistently finds that strong test-set performance does not transfer to generators the detector was not built around, and that transcoding and editing erode performance further. An acoustic engine like ours is even more exposed to that problem than a trained model.

Hybrid tracks

A human vocal over a generated backing, or a generated stem inside a human arrangement, produces exactly the segment disagreement that triggers an inconclusive result. That is appropriate: the honest answer for a hybrid track is “partly”, which a single probability cannot express.

False positives and false negatives

A false positive can damage a real musician’s reputation; a false negative merely misses a synthetic track. Those costs are not symmetrical, which is why the classification thresholds are set conservatively and why Likely AI-generated requires both a high score and agreeing segments.

How results should be used

  • As one input among several, never as the deciding factor
  • Alongside provenance: session files, stems, release history, artist conversation
  • Never as evidence in a copyright, employment, academic or enforcement decision
  • Never as the basis for a public accusation

See the detection-results disclaimer for the formal statement, or go back and analyse a track with the free AI music detector.