Riffusion
Early Riffusion generated pictures of sound. The artefacts of that approach are mostly gone now, which is itself a lesson about detection.
- Origin
- Spectrogram-image diffusion
- Current form
- Full song generation
- Detection angle
- High-frequency texture, historically banding
- Attribution supported
- No
Why old detection advice stops working
Generating audio by diffusing a spectrogram image and then inverting it produced visible horizontal banding and phase reconstruction artefacts. For a while, that was a nearly free detection signal. Newer versions do not synthesise that way, and the signal is largely gone.
This is the clearest available demonstration of a general rule: any detector tuned to a specific artefact has a shelf life. Ours is no exception, which is why the report emphasises confidence and why the accuracy page refuses to publish a single headline number.
What is measurable today
The remaining signals are the generic ones: mastering uniformity, spectral behaviour over time, stereo field movement and cross-segment agreement. None of them are Riffusion-specific, and the engine does not pretend they are.
What to listen for
Riffusion now produces full songs with vocals, so the cues worth your attention are performance cues rather than architecture artefacts.
- Consonants that arrive with the same shape every time — a sung line usually varies its attack across a verse.
- Lyric phrasing that lands exactly on the grid across every repeat of a hook.
- Section joins where the whole mix changes character at once, rather than instruments entering at slightly different moments.
- Backing vocals that sit in the identical stereo position and reverb as the lead, as if printed together.
- A room that never changes: the same ambience on the intro, the bridge and the outro.
Version drift
Riffusion is the clearest case on this site of a detection signal expiring. Spectrogram banding and phase-reconstruction smear were reliable on the original image-diffusion approach and are effectively absent from current output. Anything written before that architecture change is now actively misleading.
What survived the transition are the generic measurements — uniformity, spectral behaviour over time, stereo movement, cross-segment agreement — because they describe how production works rather than how one model works.
What editing does to the reading
The same erosion applies here as everywhere: every processing step between generation and the file you analyse removes evidence.
- Re-mastering compresses the crest-factor and micro-dynamic differences the analysis reads.
- Re-encoding moved our readings about 2.4 percentage points on its own.
- Trimming to a hook removes cross-segment agreement, and excerpt choice alone accounted for about 4.4 percentage points of movement in testing.
- Replacing a generated vocal with a recorded one genuinely changes what the file is, and a mid-range reading is the honest result.