Skip to content

Guide

Song Structure, Repetition and AI Music: What Loops Prove

"It just loops. Same eight bars over and over. Obviously AI." Repetition is the oldest device in recorded music and the most confidently misused piece of evidence in the AI argument. This is what a structure measurement actually computes — repeating cycles, near-identical passages, overall uniformity, section boundaries — what each of those genuinely tells you about how a track was assembled, and why assembly is not the same question as authorship.

· 12 min read

What a structure measurement actually computes

The tool on this site does something deliberately simple. It decodes your file in your own browser, reduces the audio to a sequence of coarse spectral fingerprints — twenty-four log-spaced frequency bands, four times a second, with the overall loudness removed so a volume change does not look like a content change — and then compares every fingerprint with every other one.

From that comparison come four readings. The strongest repeating cycle is the time offset at which the track most resembles a delayed copy of itself: usually a bar group, a loop length, or the gap between one chorus and the next. Near-identical passages count runs of two seconds or longer where a stretch of audio matches an earlier stretch almost exactly. Overall uniformity is the average similarity across the whole excerpt, a single number for how much the tonal picture changes from start to end. Section boundaries are the moments where the preceding few seconds and the following few seconds look most unlike each other.

None of that requires a model, a training set, or a guess about provenance. It is arithmetic on a spectrogram, and it describes shape. That limitation is the point: the reading is reproducible, checkable, and does not pretend to know anything it cannot see.

Repeating cycles: the loop length hiding in the numbers

When the tool reports a strongest repeat every eight seconds, it is usually telling you the tempo and the phrase length at once. At 120 BPM, a four-bar phrase in common time takes exactly eight seconds. A reported cycle of four seconds at that tempo is two bars; sixteen seconds is eight bars. If you know one of tempo or phrase length you can infer the other, which makes this the fastest way to check whether a track is built on two-bar, four-bar or eight-bar units.

The match strength alongside it matters as much as the length. A cycle at 88% means the section returns recognisably but with variation — new fills, a different vocal, an extra layer. A cycle at 99% means the return is very nearly the same audio. The first describes an arrangement. The second describes a copy.

No detected cycle is equally informative. Through-composed material, free improvisation, live recordings that drift in tempo, and orchestral writing that never restates a passage identically will all report no strong repeating structure — not because they lack repetition, but because their repetitions are performed rather than duplicated, and performance never lines up sample for sample.

Near-identical passages describe assembly, not authorship

This is the reading people most want to over-interpret. A passage that matches an earlier one above roughly 99.5% similarity across two seconds or more is, in practical terms, the same audio appearing twice. There are only a few ways that happens: someone copied a region in a DAW, a sampler or loop player fired the same file again, a pattern was rendered once and repeated, or a stem was duplicated with only a level change.

Every one of those is completely routine in human production. Duplicating a verse and editing the vocal over it is standard workflow. Loop-based genres are constructed almost entirely from repeated blocks by design. Library and background music is often deliberately built from exact repeats so it can be cut to length. Copy-and-paste is a technique, not a confession.

What the reading is genuinely useful for is the opposite direction. If a file is presented as a live take, a solo performance or a single continuous recording, exact repeats are inconsistent with that claim. A live band cannot play two seconds identically twice. That is a real finding about a specific claim — and it is far narrower and far more defensible than "this is AI".

Uniformity, and the honest version of the AI argument

Overall uniformity is where the AI question actually has some purchase, and it deserves stating carefully. Some generated output does come back unusually uniform: a tonal balance that barely moves across the whole track, because the audio was produced in a single pass as a finished stereo mix rather than assembled from parts that were recorded, arranged and mixed at different times with different decisions behind them.

The trouble is what else lands in the same range. Loop-based electronic music is uniform because it is meant to be. Background, corporate and library tracks are uniform because a bed that changes character is useless under dialogue. Anything mastered aggressively is more uniform than it was before mastering, because compression and limiting reduce exactly the variation this measure looks at. And a two-minute excerpt from the middle of a long track will read as uniform simply because it never leaves one section.

So a very high uniformity figure is a reason to keep looking, not a verdict. It says the tonal picture does not move much. Three quite different origins produce that shape, and the measurement cannot distinguish between them.

Section boundaries and edit points

Boundaries are detected as local peaks in a novelty curve: points where the spectral character of the previous couple of seconds differs most from the following couple. In practice they land on the obvious architecture of a song — the drop into a chorus, the drums cutting out, a breakdown starting, a key or instrumentation change — with a minimum spacing so a single transition is not reported four times.

Two failure modes are worth knowing. Gradual arrangements produce few or no boundaries even though the track clearly develops, because the change is spread out rather than abrupt. And a hard edit, a stitched-together file, or a sudden level jump will produce a boundary that has nothing to do with musical structure. That second case is often the more interesting one: if the boundaries do not correspond to anything you can hear as a section change, something has been cut and rejoined.

Use the trace alongside the boundary list. A novelty curve with a few tall isolated spikes describes a track with clear sections. A flat curve with no peaks describes something continuous. A curve that is noisy everywhere usually means the excerpt is too short or the material too dense for section detection to say anything useful.

How to use structure alongside a detector result

The productive framing is that structure explains a score rather than supporting it. If a detector returns a high AI-likelihood and the structure reading shows near-total uniformity, no section boundaries and several exact repeats, you have learned why the score came out that way: the file has the shape that pushes those features up. That is worth knowing precisely because it identifies a false-positive cluster — loop-based and library music will do this all day.

If a detector returns a high score and the structure reading shows genuine variety, distinct sections and only performed repetition, the score is resting on something else entirely — spectral detail, stereo behaviour, codec history — and you should go and look at those directly.

And if you are arguing about a specific claim rather than a general vibe, structure is at its strongest when it contradicts a stated provenance: exact repeats in a supposed live take, a machine-perfect cycle length in a supposedly free performance, boundaries where no edit was declared. Narrow claims that a measurement can actually support are worth more than broad ones it cannot.

What this measurement cannot do

It cannot name a generator. Nothing in a similarity matrix is specific to Suno, Udio or anything else, and any tool that claims otherwise from structure alone is guessing.

It is bounded. Very long files are measured on a central excerpt rather than end to end, so a structural feature that only occurs in the first thirty seconds of a ten-minute mix may not appear in the reading. Short files — under about thirty seconds — often cannot contain a full cycle, which is why they so frequently report no strong repeating structure.

It is easy to defeat deliberately. Adding a small amount of noise, applying a subtle random modulation, or nudging the level of every repeated section breaks exact-match detection while leaving the audio audibly unchanged. Anyone motivated to hide copy-and-paste can, in a few minutes. That places a firm ceiling on how much weight structure should carry in any dispute.

The short version

Structure measurement tells you how a track was assembled: the loop length it is built on, whether passages were copied rather than performed, how much the tonal picture moves, and where the section edges fall. All of that is genuinely useful — for reading a tempo, catching a stitched edit, or testing a claim that something was recorded live in one pass. None of it identifies a generator, and repetition on its own is a production technique rather than evidence of origin. Use it to explain why a detector scored a file the way it did, treat extreme uniformity as a false-positive risk rather than a confirmation, and keep provenance questions with metadata, stability across excerpts, and the session files.

Try the free AI music detector

Frequently asked questions

  • No. Repetition is a compositional and production choice that predates generative audio by a century. Loop-based electronic music, library beds and most pop choruses are built on exact or near-exact repeats. Repetition describes how audio was assembled, not who or what assembled it.

More reading