Skip to content

Generator hub

AI music generators, and what each one means for detection

Every generator builds audio differently, and those differences decide which measurements carry information. These pages describe how each platform works and how much a detection reading on its output is worth — including when the honest answer is “not much”.

One thing they all have in common: this tool does not perform generator attribution. It estimates whether a recording looks machine-generated, not which service produced it.

Analyse a track

Platforms covered

Side by side: what each generator produces, and what can be measured

The right-hand column is the honest part. It says which measurement carries information for that platform’s output — not that a file can be traced back to it.

Swipe the table sideways to see all columns.

GeneratorWhat it producesWhere the signal isAttribution
SunoComplete song: vocals, lyrics, arrangement, mixLoudness-maximised MP3 or WAVNo
UdioSong sections, extendable into full arrangementsCross-segment consistency and seamsNo
ElevenLabs MusicVocal-forward music generationHigh-band behaviour, indirectlyNo
Stable AudioInstrumental beds, loops, sound designSpectral ceiling and high-band textureNo
RiffusionFull song generationHigh-frequency texture, historically bandingNo
MubertRoyalty-free background and production musicSegment-to-segment repeatabilityNo
Seed MusicVocals, lyrics and instrumental in one passCross-segment tonal stabilityNo
MiniMaxComplete songs from prompts or reference clipsWeakened dynamics features; spectral evidence onlyNo
Mureka (Sonauto)Prompt-to-track with vocal controlSpectral-centroid variabilityNo
SoundrawInstrumental, royalty-free; no generated lead vocalGrid-exact timing and section-level repetitionNo

Notes on the remaining platforms

These five have shorter write-ups, so they are published here in full rather than as separate pages.

Riffusion

Early Riffusion generated pictures of sound. The artefacts of that approach are mostly gone now, which is itself a lesson about detection.

Origin
Spectrogram-image diffusion
Current form
Full song generation
Detection angle
High-frequency texture, historically banding
Attribution supported
No

Why old detection advice stops working

Generating audio by diffusing a spectrogram image and then inverting it produced visible horizontal banding and phase reconstruction artefacts. For a while, that was a nearly free detection signal. Newer versions do not synthesise that way, and the signal is largely gone.

This is the clearest available demonstration of a general rule: any detector tuned to a specific artefact has a shelf life. Ours is no exception, which is why the report emphasises confidence and why the accuracy page refuses to publish a single headline number.

What is measurable today

The remaining signals are the generic ones: mastering uniformity, spectral behaviour over time, stereo field movement and cross-segment agreement. None of them are Riffusion-specific, and the engine does not pretend they are.

What to listen for

Riffusion now produces full songs with vocals, so the cues worth your attention are performance cues rather than architecture artefacts.

  • Consonants that arrive with the same shape every time — a sung line usually varies its attack across a verse.
  • Lyric phrasing that lands exactly on the grid across every repeat of a hook.
  • Section joins where the whole mix changes character at once, rather than instruments entering at slightly different moments.
  • Backing vocals that sit in the identical stereo position and reverb as the lead, as if printed together.
  • A room that never changes: the same ambience on the intro, the bridge and the outro.

Version drift

Riffusion is the clearest case on this site of a detection signal expiring. Spectrogram banding and phase-reconstruction smear were reliable on the original image-diffusion approach and are effectively absent from current output. Anything written before that architecture change is now actively misleading.

What survived the transition are the generic measurements — uniformity, spectral behaviour over time, stereo movement, cross-segment agreement — because they describe how production works rather than how one model works.

What editing does to the reading

The same erosion applies here as everywhere: every processing step between generation and the file you analyse removes evidence.

  • Re-mastering compresses the crest-factor and micro-dynamic differences the analysis reads.
  • Re-encoding moved our readings about 2.4 percentage points on its own.
  • Trimming to a hook removes cross-segment agreement, and excerpt choice alone accounted for about 4.4 percentage points of movement in testing.
  • Replacing a generated vocal with a recorded one genuinely changes what the file is, and a mid-range reading is the honest result.

Mubert

Loop-based construction produces the most repeatable segment measurements of any category we look at.

Output
Royalty-free background and production music
Construction
Generative arrangement of loop material
Detection angle
Segment-to-segment repeatability
Attribution supported
No

Repetition as a measurement

When a track is assembled from recombined loops, consecutive segments can be almost statistically identical: the same spectral centroid, the same crest factor, the same stereo distribution. Live performance never does this, because a drummer hits differently every bar.

But so does a lot of human music

Programmed electronic music, library music cut to a grid, and anything built on a two-bar loop in a DAW behave the same way. Repeatability is evidence of a machine somewhere in the chain — it is not evidence that the machine composed the music. The engine treats it as a moderate-weight feature and the report says so in plain language.

What to listen for

Mubert-style production music is designed to sit under something else, so the cues are structural rather than performative.

  • Arrangements that add and remove layers without ever developing a melodic idea.
  • Transitions that are pure layer swaps: nothing anticipates the change, nothing resolves after it.
  • Percussion that repeats bar for bar with no ghost notes, fills or timing drift.
  • An ending that fades or stops on the grid rather than being played to a close.
  • The same reverb tail on every element, because the loops were mixed into one shared space.

Version drift

Loop-based generative services change their libraries and their arrangement logic far more often than they change architecture, so the acoustic character of output moves gradually rather than in jumps. Newer material tends to include more recorded instrument content, which softens the extreme segment repeatability older output showed.

That direction of travel matters: repeatability is a signal that is weakening over time, not strengthening. Do not treat a merely tidy arrangement as decisive.

What editing does to the reading

Background music is usually edited before anyone hears it, which is exactly the processing that erodes the reading.

  • Cutting a track to length for a video removes the cross-segment comparison the repeatability feature depends on.
  • Ducking under a voiceover changes the dynamic profile the engine measures.
  • Re-encoding for delivery moved our readings about 2.4 percentage points on its own.
  • Excerpt choice alone accounted for about 4.4 percentage points of movement in testing, so analyse the full delivered file where you can.

Seed Music

Single-pass generation of a whole song produces a track with no production history — and that absence is measurable.

Output
Vocals, lyrics and instrumental in one pass
Lineage
Research song-generation systems
Detection angle
Cross-segment tonal stability
Attribution supported
No

No production history

A conventional record accumulates variation: different takes, different processing on different sections, a mastering engineer making a chorus brighter than a verse. An end-to-end generation has one origin, so tonal balance across the track tends to hold far steadier than a produced record does.

The engine reads this as high cross-segment agreement. On its own that is not damning — a live one-take recording is also consistent — but combined with mastering uniformity it forms a coherent pattern.

What to expect from the report

Expect strong segment agreement and a confidence level that reflects it. Expect the report to attribute the result to measurements — brightness stability, dynamic range, high-band structure — rather than to any named platform.

What to listen for

End-to-end systems sing and play at once, so the tells sit in how the performance relates to itself rather than in any single sound.

  • A vocal that never moves off-mic: no leaning in on a loud line, no backing off on a soft one.
  • Breaths placed at plausible moments but with the same length and level every time.
  • Instruments that never mask each other awkwardly, because they were generated as one balanced whole.
  • A bridge that sounds like the verse with different notes rather than a different recording decision.
  • Ambience that is identical in the first bar and the last, with no build-up of room or tape character.

Version drift

Research-lineage systems iterate quickly and publish little. In practice the visible direction is towards more deliberate variation: newer output introduces section-to-section differences in tone and level that earlier versions did not have, precisely because uniformity was the obvious criticism.

That erodes the cross-segment feature these pages describe. Treat strong segment agreement as evidence that ages, and weight provenance more heavily as the models improve.

What editing does to the reading

Anything done after generation reduces what the analysis can see.

  • Mixing sections at different levels breaks the tonal stability the engine reads.
  • Adding a recorded instrument or a live vocal genuinely changes what the file is; a mid-range reading is then the honest answer.
  • Re-encoding shifted our readings by about 2.4 percentage points on its own.
  • Trimming to an excerpt removed the comparison entirely and accounted for about 4.4 percentage points of movement in testing.

MiniMax

When a model copies the dynamics of a human reference, the features built on dynamics stop working.

Output
Complete songs from prompts or reference clips
Distinctive feature
Reference conditioning
Detection angle
Weakened dynamics features; spectral evidence only
Attribution supported
No

Inherited dynamics

Reference conditioning lets the model take the loudness envelope and general character of an existing recording. If the reference is a human track with natural crest factor, the generated output can inherit that profile — and crest factor is one of the cheapest, most commonly used detection features in the field.

This is a real weakness, not a hypothetical one. The engine reduces the weight of dynamics-derived evidence when other measurements disagree with it, and reports lower confidence rather than pretending the feature still holds.

What still carries signal

Spectral structure in the upper bands, stereo field movement over time, and cross-segment agreement are less affected by reference conditioning, because they depend on how the audio was synthesised rather than on how loud it is.

What to listen for

Reference-conditioned output borrows a shape it did not earn, and that mismatch is often audible before it is measurable.

  • Dynamics that swell convincingly while the parts underneath do not change how they are played.
  • A drum performance that gets louder without getting harder — no change in attack or timbre with level.
  • Timbres that stay identical through a section the arrangement treats as a climax.
  • Transitions borrowed wholesale: a fill that fits the energy curve but not the groove around it.
  • Vocal intensity that tracks the mix level rather than the singer's effort.

Version drift

Multimodal families update their audio components alongside everything else, often without a music-specific release note. The practical consequence is that the balance between inherited and generated characteristics changes without warning between one month's output and the next.

Because of that, treat the dynamics features as unreliable for this category by default rather than only when a reference is known to have been used.

What editing does to the reading

With dynamics already compromised, the remaining spectral evidence is the part editing damages most.

  • Re-mastering alters high-band structure, which is where most of the surviving signal sits.
  • Re-encoding moved our readings about 2.4 percentage points on its own.
  • Excerpt choice accounted for about 4.4 percentage points, so analyse the full file.
  • Layering generated material under recorded parts produces a genuinely mixed file, and a mid-range reading reflects reality rather than failing.

Mureka (Sonauto)

A style-transfer pass rewrites the surface of a track, and surfaces are what spectral measurement reads.

Also known as
Sonauto
Output
Prompt-to-track with vocal control
Detection angle
Spectral-centroid variability
Attribution supported
No

Smoothed brightness

Live performance produces constant small changes in brightness: a pick attack, a cymbal, a consonant, a fret buzz. Measured over time, spectral centroid in a human recording jitters. A style-transfer pass tends to regularise that jitter, producing a flatter brightness curve than the same arrangement played by people.

What else flattens brightness

Multiband compression, aggressive limiting, spectral-balance plug-ins and AI mastering services all reduce brightness variability in human recordings. Any of them can push a genuine performance towards the same reading, which is why this feature is weighted moderately rather than decisively.

What to listen for

Style transfer keeps the musical gesture and replaces the sound, so listen for the seam between the two.

  • Transients that feel blunted: a snare that arrives without the crack that should precede its body.
  • Consistent brightness across sections that should differ, such as a sparse verse and a full chorus.
  • A vocal timbre that stays put while the phrasing suggests it should open up.
  • High-frequency detail that sounds painted on rather than produced by the instruments underneath.
  • Sibilance that behaves identically on every line, regardless of the word being sung.

Version drift

This service has changed name and model generation, and vocal control has improved noticeably. The consequence for detection is that flattened brightness is less pronounced than it was: newer output restores more transient detail, which moves it closer to produced human material.

Anything written about this platform more than a version ago should be re-checked before being relied on, the same caution the Riffusion page makes explicit.

What editing does to the reading

Style transfer is itself an editing step, and further edits stack on top of it.

  • Exciters and high-shelf boosts reintroduce brightness variation, pushing the reading down.
  • AI mastering pushes it back up, sometimes on a completely human recording.
  • Re-encoding shifted our readings by about 2.4 percentage points on its own.
  • Excerpt choice accounted for about 4.4 percentage points, so use the whole track and read the stability check.

Why attribution is not offered

Naming the tool behind a track requires a classifier trained on labelled output from every platform in question, kept current as each of them ships new model versions. Nobody has published such a model with credible held-out results, and we are not going to imply one exists by putting platform names on a result screen.

What these pages give you instead is context: knowing that a generator maximises loudness by default, or stitches sections together, tells you which parts of a report to weigh and which to discount.

Related reading