Research
This page will hold our evaluations. Right now it holds the plan, because publishing an invented benchmark would be worse than publishing nothing.
Last updated August 2026
Why there are no numbers here yet
A benchmark is only meaningful if the dataset, sampling procedure, generator versions and transformations are documented well enough for someone else to reproduce it. Assembling that for music requires licence-clear human recordings and generated material we are permitted to redistribute or at least describe precisely. That work is underway; the results will appear here whether or not they flatter the detector.
Planned evaluations
1. Recompression robustness
The same tracks encoded at lossless, 320, 192, 128 and 64 kbps, plus a 44.1 → 22.05 kHz resample. Question: how much of the detector’s signal survives ordinary distribution? Our expectation, stated in advance, is that the spectral-ceiling feature becomes useless below 192 kbps.
The full design, the pre-registered per-bitrate expectations and the compensations already built into the engine are written up in the compression and detection study.
2. Clip length
Identical tracks truncated to 10, 20, 30, 60 and 120 seconds. Question: at what duration does segment agreement become informative? Pre-registered expectation: little value below 30 seconds.
3. Unseen generators
Evaluation restricted to output from generators whose characteristics were not consulted when setting the engine weights. This is the test that matters most, and the one most detector marketing quietly avoids.
4. Instrumental vs vocal
Matched pairs of instrumental and vocal material. Question: how much of any measured performance depends on vocal artefacts?
5. Post-production resistance
Generated tracks after tempo change, pitch shift, EQ, re-mastering and analogue-style saturation. Question: how cheaply can the detector be defeated?
How results will be reported
- Research question, stated before the test is run
- Dataset description, sampling procedure and generator versions
- Every transformation applied, with parameters
- Confusion matrix, precision, recall, false-positive and false-negative rates
- Limitations, including tests that did not work
- Enough detail for independent reproduction
Unfavourable results will be published too. A detector site that only ever reports good news is not a research page, it is an advertisement.
Model version history
Version changes and their rationale are documented on the methodology page. The current version is the AI Music Detector analysis engine.
Collaboration
If you hold a licence-clear evaluation set, or want to run an independent test, get in touch via the contact page. Independent criticism will be published alongside our response.