Skip to content
Research

Research

AI music detection is an evolving technical problem. This section documents how we evaluate the detector, what we have measured, what remains untested, and where the method can fail.

Updated 2026-08-30 · engine build ensemble-1

Nothing on these pages is estimated. Where a figure has not been measured, the page says so rather than filling the gap. Two studies are complete; the accuracy benchmark that people usually want a single number from is not, and we would rather say that plainly than publish a percentage we cannot defend.

Measured

Completed studies. Each one states its question before its result, and each is a study we could run without a licence-clear labelled corpus.

Completed

Human electronic music false positives

Question. How often does the detector read synthesiser-only, quantised human production as AI-generated?

Finding. On twelve synthesiser-only reference renders with no generator involved, three read above 70% AI probability, and the engine declined to issue a verdict on 83 of 84 files across the encode ladder.

Completed

Compression and transcoding sensitivity

Question. How much does an MP3 encode move the score, relative to the lossless master?

Finding. Mean shift of 2.4 pp across the 320/192/128/64 kbps ladder, 0.0 pp run-to-run variation, and a 10-second excerpt moving the score further (4.4 pp mean) than any bitrate does.

Methodology

The evaluation protocol — objective, dataset design, the holdout principle, the transformation matrix and the exact definition of every metric — is published in full, before results exist, at the benchmark protocol. Publishing the method first is deliberate: it prevents us fitting an explanation to whatever the data eventually says.

How the product itself works, and what it does and does not do, is documented separately on how AI music detection works. What a result means in practice, and every failure mode we know about, is on accuracy and limitations.

Planned

Designed, not run. These entries carry no findings and will not until the work is actually done.

What are precision, recall, specificity and the false-positive rate on verified human and verified AI-generated recordings?

Does performance hold on output from a generator that was never consulted when designing thresholds?

Clip duration sensitivity

Planned

At what duration does the result become stable enough to be worth reporting?

Post-production resistance

Planned

How much ordinary mixing, mastering and re-recording is needed before generated audio reads as human?

Vocal versus instrumental material

Planned

How much of any measured performance depends on vocal artefacts?

References and further reading

  • Benchmark protocol — objective, corpus design, metric definitions, transformation matrix
  • benchmark.json — machine-readable facts, CC BY 4.0, null where unmeasured
  • Editorial standards — how these pages are written and corrected
  • Contact — corrections, replication attempts and corpus contributions