Skip to content

Stable Audio detection

Diffusion models are trained up to a bandwidth. Where that bandwidth ends is sometimes visible — and sometimes indistinguishable from a codec.

Output
Instrumental beds, loops, sound design
Method
Diffusion in a learned audio representation
Detection angle
Spectral ceiling and high-band texture
Attribution supported
No

Upload Audio or Drag & Drop

Upload an audio file to receive a probabilistic analysis of characteristics associated with AI-generated music.

  • MP3
  • WAV
  • M4A
  • FLAC
  • OGG
  • AAC
  • WebM

MP3, WAV, FLAC, AAC, M4A, OGG, WebM · max 25 MB · min 10 seconds · 30+ seconds recommended

Your audio never leaves your device. Decoding and analysis run entirely in this browser tab, and nothing is uploaded to a server. Your audio is processed only to perform this analysis, and your uploaded audio and temporary analysis data are automatically deleted after processing. No report links are created, and your analysis is never publicly accessible. Upload only audio you are authorised to process — analysis does not transfer ownership or publishing rights.

  • Free
  • Fast
  • Secure
  • No registration

Where the spectrum stops

A diffusion model generates inside the representation it was trained on. If that representation is band-limited, the output is band-limited too, and energy above the limit falls away with a sharpness that natural recordings rarely show. A cymbal recorded with a decent microphone has energy that tapers; a hard shelf at a suspiciously round frequency does not look like a cymbal.

The codec confound

Every lossy encoder also imposes a ceiling. A 128 kbps MP3 typically cuts around 16 kHz with a similar sharpness. From the file alone, a model ceiling and a codec ceiling can be functionally identical, which means the feature is only trustworthy on high-bitrate or lossless material.

The engine handles this by discounting the spectral-ceiling feature when other evidence suggests the file has been re-encoded, and by lowering confidence rather than inventing certainty. This confound is the subject of our compression study.

Instrumental material is harder

Without a vocal, several of the strongest human-performance cues disappear: breath, phrasing irregularity, the micro-timing of a sung line. Ambient and loop-based instrumentals are among the hardest categories for any detector, human or automatic, and results on them deserve extra scepticism.

Stable Audio detection FAQ

  • Because lossy encoding removes the high-frequency detail several features depend on. The lossless reading is the more reliable one; a low-bitrate copy should generally be read as lower confidence.

Other generators