Stable Audio detection
Diffusion models are trained up to a bandwidth. Where that bandwidth ends is sometimes visible — and sometimes indistinguishable from a codec.
- Output
- Instrumental beds, loops, sound design
- Method
- Diffusion in a learned audio representation
- Detection angle
- Spectral ceiling and high-band texture
- Attribution supported
- No
Upload Audio or Drag & Drop
Upload a song to check whether its vocals or instrumental content show signs of AI generation. Analysis is performed by our AI music detection system.
- MP3
- WAV
- FLAC
- AAC
- M4A
- MP4
- OGG
- OPUS
MP3, WAV, FLAC, AAC, M4A, MP4, OGG, OPUS · max 25 MB (our upload limit) · 30+ seconds recommended
Your audio is uploaded over an encrypted connection and sent to our AI music detection partner solely to perform this analysis. We do not create a public report page and we do not intentionally retain your uploaded audio after the analysis completes. Upload only audio you are authorised to process — analysis does not transfer ownership or publishing rights.
- Free
- Fast
- Secure
- No registration
Where the spectrum stops
A diffusion model generates inside the representation it was trained on. If that representation is band-limited, the output is band-limited too, and energy above the limit falls away with a sharpness that natural recordings rarely show. A cymbal recorded with a decent microphone has energy that tapers; a hard shelf at a suspiciously round frequency does not look like a cymbal.
The codec confound
Every lossy encoder also imposes a ceiling. A 128 kbps MP3 typically cuts around 16 kHz with a similar sharpness. From the file alone, a model ceiling and a codec ceiling can be functionally identical, which means the feature is only trustworthy on high-bitrate or lossless material.
The engine handles this by discounting the spectral-ceiling feature when other evidence suggests the file has been re-encoded, and by lowering confidence rather than inventing certainty. This confound is the subject of our compression study.
Instrumental material is harder
Without a vocal, several of the strongest human-performance cues disappear: breath, phrasing irregularity, the micro-timing of a sung line. Ambient and loop-based instrumentals are among the hardest categories for any detector, human or automatic, and results on them deserve extra scepticism.
What to listen for before you measure anything
Diffusion instrumentals rarely fail in an obvious way. They fail by being too even. Put the file on decent headphones and listen for the following, then treat the number as a second opinion rather than the first.
- Reverb tails that behave identically on every hit, as if one space were printed once and reused rather than played in.
- Percussion transients that all land with the same attack shape — no rim of a stick catching the edge, no accent that overshoots.
- A high band that sounds smooth but featureless: air without the grain of a real cymbal or a room.
- Loops that repeat without the tiny performance drift a played part accumulates over four or eight bars.
- Endings that fade rather than resolve, which is a strong hint the clip was generated to a duration rather than written to one.
Version drift
Successive Stable Audio releases have raised the training bandwidth and improved stereo behaviour. The practical effect is that the sharp spectral shelf which made early output easy to spot is far less dependable on recent material, while the evenness cues above have barely moved. Advice written against an older release will over-trust the ceiling feature and under-weight uniformity.
This is why the report never quotes a per-generator accuracy figure. A feature that worked on last year's output is not a feature you can price today.
What editing does to the reading
Diffusion beds are usually the raw material for something else — layered under a vocal, chopped into a loop, re-mastered inside a project. Each of those steps removes evidence.
- Layering live or sampled parts over the bed reintroduces genuine performance variance and pulls the reading down, correctly, because the finished work is now partly human.
- Re-mastering with a limiter flattens the crest-factor difference the analysis reads.
- Re-encoding imposes its own ceiling; our compression study measured about 2.4 percentage points of movement from encoding alone.
- Chopping to a short loop removes cross-segment agreement entirely, which is the single most reliable feature on instrumental material.
Stable Audio detection FAQ
Because lossy encoding removes the high-frequency detail several features depend on. The lossless reading is the more reliable one; a low-bitrate copy should generally be read as lower confidence. Re-encoding alone moved readings about 2.4 percentage points in our robustness testing.