Detect AI-generated music
Detection is a measurement problem, not a lookup. This page explains which acoustic signals carry information, why each one is individually weak, and what combining them actually buys you.
Upload Audio or Drag & Drop
Upload an audio file to receive a probabilistic analysis of characteristics associated with AI-generated music.
- MP3
- WAV
- M4A
- FLAC
- OGG
- AAC
- WebM
MP3, WAV, FLAC, AAC, M4A, OGG, WebM · max 25 MB · min 10 seconds · 30+ seconds recommended
Your audio never leaves your device. Decoding and analysis run entirely in this browser tab, and nothing is uploaded to a server. Your audio is processed only to perform this analysis, and your uploaded audio and temporary analysis data are automatically deleted after processing. No report links are created, and your analysis is never publicly accessible. Upload only audio you are authorised to process — analysis does not transfer ownership or publishing rights.
- Free
- Fast
- Secure
- No registration
Why there is no single tell
Early generation systems left obvious traces: smeared transients, a hard frequency ceiling, metallic reverb tails, repeating micro-textures. Those were easy to detect and, consequently, the first things generator developers fixed. What remains is statistical: distributions of energy, movement and structure that differ slightly, on average, between generated and performed recordings.
“Slightly, on average” is the crucial phrase. It means detection works better across many tracks than on any single one — the opposite of what a user wants.
The signal families
Spectral artefacts
Where a recording’s energy stops, how sharply it stops, and how the noise floor behaves above that point. Strongly confounded by lossy encoding: the same measurement fires for a 128 kbps MP3 of a jazz quartet.
Structural regularity
Generated arrangements often hold a very uniform tonal balance across sections, because they are produced in one pass rather than tracked, edited and mixed. Comparing the spectra of several segments captures this, and it is currently the most useful signal we measure. It is also destroyed the moment someone remixes the track.
Micro-dynamics
Crest factor and frame-to-frame timbral movement describe how much a recording breathes. Useful in acoustic genres; nearly worthless in loudness-maximised electronic music, where human masters look identical.
Learned embeddings
A model trained on labelled audio can pick up patterns no hand-written rule would find. The catch, reported consistently in the research literature, is that such models generalise poorly to generators they never saw, and degrade under transcoding. They are powerful and brittle at the same time.
Provenance and watermarks
The only category that offers real certainty. Where a generator embeds a watermark, or a file carries Content Credentials, you get cryptographic evidence rather than a statistical guess. Coverage is patchy and metadata is easily stripped, but when present this beats every acoustic method combined.
What an ensemble adds
Combining weak signals helps only when their errors are uncorrelated. Spectral ceiling and high-band energy largely measure the same physical thing, so stacking them adds little; structural regularity and micro-dynamics fail in different situations, so combining those genuinely helps. Our current weighting reflects that, and is documented openly in the methodology.
What defeats detection
- Re-recording the output through speakers and a microphone
- Re-mixing generated stems with human-played parts
- Tempo, pitch and EQ changes followed by fresh mastering
- Aggressive lossy transcoding
- Simply using a newer generator than the detector was built around
None of these are secrets, and pretending otherwise would give a false sense of security. That is exactly why we publish what the detector cannot do alongside the tool.