Skip to content

Guide

Tempo, Timing Grids and AI Music: What Quantisation Proves

"The drums are inhumanly perfect. No player is that tight. It has to be AI." Timing is the most confidently misread measurement in the whole argument, because the thing it detects — a machine clock — became standard studio equipment forty years before generative audio existed. This is what a tempo and timing reading actually computes, what grid-tight placement genuinely proves, and the narrow set of claims timing can settle properly.

· 12 min read

What a timing measurement actually computes

The tool on this site decodes your file in your own browser and reduces it to an onset envelope: a curve describing how much new spectral energy appears from one short frame to the next. Hits, plucks, consonants and transients push that curve up; sustained notes and pads leave it flat. Peaks in the curve are taken as onsets — the moments something starts.

Tempo comes from autocorrelating that envelope, which is a formal way of asking at what delay the curve most resembles itself. A steady pulse produces a strong peak at the beat period, and the strongest period between roughly 60 and 200 BPM becomes the reported tempo.

Given a tempo, the tool builds a sixteenth-note grid and searches for the phase — the grid's alignment in time — that best fits your actual onsets, rather than assuming the file starts on a beat. It then reports the typical distance from each onset to the nearest grid line, the share of onsets landing within a few milliseconds of one, how widely those distances spread, and whether the first and second halves of the excerpt fit different tempi.

There is no model and no training set anywhere in that chain. It is arithmetic on a spectrogram, which is exactly why the numbers are reproducible and exactly why they cannot answer a provenance question by themselves.

Quantisation is older than generative audio by decades

Drum machines have placed hits on an exact clock since the late 1970s. Step sequencers have done the same since before that. Every DAW since the early 1990s has shipped with a quantise command that drags a recorded hit to the nearest grid line, and using it is unremarkable craft rather than a shortcut. Entire genres — house, techno, trap, drum and bass, most modern pop production — are defined by placement a human cannot achieve by hand.

So a track measuring half a millisecond from the grid has told you one thing clearly: its rhythm was produced by a machine clock. That covers a generated render, a programmed pattern, a quantised live take, and a loop pack. The measurement does not distinguish between them because nothing in the placement itself differs.

The inverse claim fails just as hard. Models trained on performed music reproduce performed timing, and a great deal of generated output shows the same small, human-looking deviations a session player leaves behind. Loose timing is not evidence of a human, and tight timing is not evidence of a machine author. Both directions of the inference are unsound.

Reading the numbers without over-reading them

Median off-grid distance is the headline. Under about 3 ms means placement is effectively exact — a sequencer, a render, or a heavily quantised edit. Between roughly 5 and 15 ms is the band that tight players and lightly corrected takes occupy: audible as precise, measurably human. Beyond about 25 ms you are looking at either genuinely loose playing or a grid model that does not describe the feel.

On-grid share and spread must be read together, because they answer different questions. A high share with a near-zero spread means uniform mechanical placement: almost everything is on a line, and the small residue is measurement noise. A high share with a wide spread usually means a hybrid — a programmed backbone holding the grid while performed parts float around it, which is how most commercial records are actually built.

Tempo drift compares the two halves of the measured excerpt. A meaningful difference suggests either a performance recorded without a click, a deliberate tempo change, or a stitched edit between two files. Zero drift across four minutes is a machine clock, and in a file presented as a live take that absence is as informative as the placement.

Why the tempo sometimes reads half or double

Autocorrelation finds periodicity, and music is periodic at several metrical levels at once. A track at 140 BPM with strong eighth notes is genuinely self-similar at 280 BPM and at 70 BPM, so the estimator can settle on any of them. This is a known and well-documented property of every tempo estimator, not a bug in this one.

A perceptual prior helps — human tapping clusters around 120 BPM, so candidates near that range are weighted up — but it cannot resolve every case. If the reported figure is half or double what you tap, halve or double it and move on.

Importantly, the grid measurements survive that ambiguity. Sixteenth-note spacing scales with the tempo, so a doubled tempo estimate simply produces a grid twice as fine, and onsets that sat on lines still sit on lines. The off-grid figures remain meaningful even when the BPM label is at the wrong metrical level.

The claims timing genuinely settles

Move from "is this AI" to "is this consistent with what I was told", and timing becomes a strong tool rather than a weak one.

A file presented as a live band take, measuring machine-exact placement with no tempo drift across four minutes, is inconsistent with that description. Humans do not hold a clock to a millisecond for that long, and the total absence of drift is the harder finding of the two. That is a checkable statement about a specific claim.

A file presented as one continuous recording, whose tempo steps cleanly between halves, points to an edit or a splice. A file presented as a solo acoustic performance whose onsets sit uniformly on a sixteenth grid points to programming or heavy correction. In each case the conclusion is about production history, not authorship — and production history is exactly the kind of thing that holds up when someone pushes back.

Where timing readings mislead

The grid model here is straight sixteenth-note division. Swung feels, shuffled hi-hats, triplet subdivisions and any deliberate push or pull against the beat will report large deviations that are entirely intentional. A jazz drummer playing beautifully behind the beat measures as loose because the measurement has no concept of behind.

Onset detection needs transients. Ambient, orchestral, choral, drone and solo-voice material produce a smooth envelope with few clear peaks, and heavy reverb smears whatever peaks exist. The tool reports that it cannot measure rather than inventing numbers, which is the right behaviour and also a real coverage limit.

Lossy encoding moves things slightly. A browser reconstructing a low-bitrate MP3 places transients a millisecond or two from where they sat in the master, which matters when you are reading a 3 ms threshold. And drift is compared between two halves only, so a slow continuous slide can appear as a small step rather than the gradual movement it is.

Using timing alongside everything else

Timing is context for a detector score, not a replacement for one. When a scan lands in the middle and you want to know why, machine-exact placement tells you the file was sequenced, which is one of the strongest false-positive clusters in the entire field: electronic and programmed music scores higher on average for reasons that have nothing to do with origin.

Read it with the other free measurements here. Structure and repetition show whether passages were copied rather than performed. File metadata carries encoder and editor history. A spectrogram shows band structure and codec history. Stereo field shows mix decisions. None of these is a verdict, and together they describe a production history that either matches a stated provenance or does not.

If the question genuinely matters — a dispute, a rights claim, a submission review — the decisive evidence is almost never in the audio. It is the session files, the project history, the timestamps and the person who can produce them. Timing measurement helps you ask better questions of that material; it does not substitute for it.

The short version

A timing reading tells you whether a rhythm was placed by a clock or by hands, how tightly it holds, and whether the tempo moves. That is genuinely useful information: it explains why sequenced music scores higher on detectors, it flags stitched edits, and it can test a specific claim that something was recorded live in one pass. What it cannot do is establish origin, because quantisation is forty-year-old craft and generated output frequently imitates performed timing. Treat grid-perfect placement as evidence of sequencing and as a false-positive risk, read it beside structure, metadata, spectrogram and stereo findings, and keep the provenance question with the session files.

Try the free AI music detector

Frequently asked questions

  • No. Grid-exact placement means the rhythm came from a machine clock — a drum machine, a step sequencer, or a quantised take — and that has been normal studio practice since the late 1970s. It covers most electronic music and most modern pop production, so it identifies sequencing, not authorship.

More reading