Stable Audio detection
Diffusion models are trained up to a bandwidth. Where that bandwidth ends is sometimes visible — and sometimes indistinguishable from a codec.
- Output
- Instrumental beds, loops, sound design
- Method
- Diffusion in a learned audio representation
- Detection angle
- Spectral ceiling and high-band texture
- Attribution supported
- No
Upload Audio or Drag & Drop
Upload an audio file to receive a probabilistic analysis of characteristics associated with AI-generated music.
- MP3
- WAV
- M4A
- FLAC
- OGG
- AAC
- WebM
MP3, WAV, FLAC, AAC, M4A, OGG, WebM · max 25 MB · min 10 seconds · 30+ seconds recommended
Your audio never leaves your device. Decoding and analysis run entirely in this browser tab, and nothing is uploaded to a server. Your audio is processed only to perform this analysis, and your uploaded audio and temporary analysis data are automatically deleted after processing. No report links are created, and your analysis is never publicly accessible. Upload only audio you are authorised to process — analysis does not transfer ownership or publishing rights.
- Free
- Fast
- Secure
- No registration
Where the spectrum stops
A diffusion model generates inside the representation it was trained on. If that representation is band-limited, the output is band-limited too, and energy above the limit falls away with a sharpness that natural recordings rarely show. A cymbal recorded with a decent microphone has energy that tapers; a hard shelf at a suspiciously round frequency does not look like a cymbal.
The codec confound
Every lossy encoder also imposes a ceiling. A 128 kbps MP3 typically cuts around 16 kHz with a similar sharpness. From the file alone, a model ceiling and a codec ceiling can be functionally identical, which means the feature is only trustworthy on high-bitrate or lossless material.
The engine handles this by discounting the spectral-ceiling feature when other evidence suggests the file has been re-encoded, and by lowering confidence rather than inventing certainty. This confound is the subject of our compression study.
Instrumental material is harder
Without a vocal, several of the strongest human-performance cues disappear: breath, phrasing irregularity, the micro-timing of a sung line. Ambient and loop-based instrumentals are among the hardest categories for any detector, human or automatic, and results on them deserve extra scepticism.
Stable Audio detection FAQ
Because lossy encoding removes the high-frequency detail several features depend on. The lossless reading is the more reliable one; a low-bitrate copy should generally be read as lower confidence.