Guide
Pitch, Tuning and AI Music: What Perfect Intonation Proves
"The vocal is inhumanly in tune. No singer holds a note that flat-line steady. It has to be AI." Pitch is the measurement people trust most and misread fastest, because the thing it detects — mathematically exact intonation — has been a one-click studio feature since the late 1990s. This is what a pitch and tuning reading actually computes, what grid-exact intonation genuinely proves, and the narrow set of claims pitch can settle properly.
· 12 min read
What a pitch measurement actually computes
The tool on this site decodes your file in your own browser, mixes it to mono, and walks through it in short overlapping frames. In each frame it asks a single question: what delay makes this slice of waveform most resemble itself? That is autocorrelation, and for a periodic sound the answer is the period of the fundamental. One over that period is the frequency, and the frequency converts to a note plus an offset in cents.
Frames where the self-similarity peak is weak — percussion, noise, breath, silence, heavy distortion — are discarded rather than guessed at. What survives is a pitch track: a list of times and frequencies for the parts of your file that were confidently periodic. Everything reported afterwards is arithmetic on that list.
Reference tuning comes next. Instead of assuming A440, the tool tests every semitone grid from 50 cents flat to 50 cents sharp and keeps whichever grid the tracked pitches collectively fit best. That single step is what makes the rest of the numbers trustworthy: a file mastered at 435 Hz would otherwise look wildly out of tune when it is merely tuned somewhere else.
Against that best-fit grid the tool reports the median distance from each tracked pitch to the nearest semitone in cents, the share of frames landing within 15 cents of a semitone, and the median frame-to-frame pitch movement inside continuous voiced runs — the jitter figure. There is no model, no training set and no classifier anywhere in that chain, which is exactly why the numbers reproduce and exactly why they cannot answer a provenance question alone.
Cents, and how much of one a human can hear
A cent is one hundredth of an equal-tempered semitone, so an octave holds 1,200 of them. Trained listeners resolve pitch differences of roughly 5 to 10 cents on sustained tones in isolation, and considerably worse inside a dense mix where masking hides fine detail. This matters because the interesting thresholds in pitch analysis sit right at the edge of audibility.
Under about 6 cents median deviation, intonation is effectively exact — closer than a human sustains without help. Between 6 and 20 cents is the ordinary band for competent singing and playing, including most professionally released vocals after light correction. Above 20 cents you are looking at expressive movement, portamento and vibrato passing through the measurement, non-equal tuning systems, or a mix so dense that the tracker is following a blend rather than a note.
Those numbers describe the median, not the maximum, and that distinction is important. A vocal can sit at 8 cents median while containing individual bends of 80 cents, because a scoop into a note or a blues bend out of one is deliberate and brief. Reading the median as a ceiling is one of the commonest errors people make with this tool.
Pitch correction is older than generative audio by a generation
Antares released Auto-Tune in 1997. Melodyne's note-level pitch editing arrived in 2001. Both spread through commercial production immediately, and by the mid-2000s tuning a lead vocal was as unremarkable as compressing it. The graphical mode of those tools does not just move a note to the nearest scale degree — on aggressive settings it flattens the drift and vibrato inside the note as well, producing precisely the still, grid-locked pitch line that people now read as machine authorship.
Everything else in a modern arrangement is exact by construction. A software synthesiser plays the note it is asked for, to the cent. A sampler transposes by exactly the interval requested. MIDI-triggered instruments have no intonation error to begin with. So a file measuring two cents from the grid has told you one thing clearly: software determined the pitch. That covers a generated render, a programmed synth part, a corrected vocal and a sample library equally, because nothing in the intonation itself differs between them.
The inverse claim collapses just as quickly. Models trained on performed music reproduce performed intonation: generated vocals routinely carry scoops, vibrato, drift and the small errors singers leave behind, because those patterns were in the training data and are part of what makes the output convincing. Imperfect pitch is not evidence of a human. Perfect pitch is not evidence of a machine author. Neither direction of that inference survives contact with how records are actually made.
Reading the four numbers together
Median deviation and in-tune share answer nearly the same question from opposite ends, so they usually agree; when they disagree, the disagreement is the finding. A high in-tune share with a large median means most frames are locked but a minority sit far off — typically a corrected vocal with uncorrected harmonies underneath, or a tuned lead over a detuned pad.
Jitter is where the more interesting reading lives. Jitter measures how much pitch moves from one frame to the next inside a sustained run, and it separates two things the deviation figure cannot. Near-zero jitter with a high in-tune share is a mathematically still pitch line: synthesis, or hard correction with the drift removed. A few cents of jitter with the same high in-tune share is a corrected performance that still breathes — a real singer moved onto the grid without having the life ironed out. Human sustained notes without correction typically show a handful of cents of jitter plus vibrato at roughly 4 to 7 Hz.
Reference tuning is the quiet one that occasionally matters most. A result within a few cents of 440 Hz tells you nothing at all, because that is the modern default. A clear offset is genuinely informative: it appears in period instruments and alternative tunings, in transfers from tape or vinyl running slightly off speed, in files converted through a sample-rate mismatch, and in re-uploads that were deliberately pitch-shifted to slip past content matching. That last case is the single most practically useful thing this tool surfaces, and it has nothing to do with AI.
The claims pitch genuinely settles
Move the question from "is this AI" to "is this consistent with what I was told", and pitch becomes a strong instrument rather than a weak one.
A file presented as a raw, unprocessed live vocal, whose pitch snaps to semitones within a couple of cents and shows no drift inside sustained notes, is inconsistent with that description. Something processed it. That is a checkable statement about a specific claim, and it holds up when someone pushes back, because the alternative explanations — correction, synthesis, sample replacement — all still contradict "unprocessed".
A recording presented as an original transfer whose best-fit tuning sits 30 cents off concert pitch was either played at that tuning, captured at the wrong speed, or shifted after the fact. A stem presented as a live instrument that measures cent-exact across every note was played by software. And a track suspected of being a re-upload of someone else's master will often reveal the shift here before anything else in your workflow catches it.
In every one of those cases the conclusion is about production history, not authorship. Production history is the kind of claim that survives scrutiny; authorship claimed from audio alone is not.
Where pitch readings mislead
The tracker follows one fundamental at a time. In a dense mix it locks onto whatever periodic component dominates each frame, which can switch between bass, lead vocal and pad within a second. The arithmetic stays correct while describing something you never intended to measure. Solo voice, exposed lead lines and sparse arrangements are where these numbers mean what you think they mean.
Equal temperament is assumed. Just intonation, meantone, maqam and other microtonal systems will register as substantial deviation, as will a genuinely well-tuned barbershop chord, because pure intervals are not equal-tempered intervals. Slide guitar, fretless bass, blues vocal phrasing, gliding synth leads and heavy portamento all report large medians for entirely intentional reasons.
Octave errors happen. Autocorrelation can settle on a period twice or half the true one, especially where a strong second harmonic dominates or the fundamental is weak. The reported note may be an octave out while the cents offset within the octave stays correct, so a suspicious octave does not invalidate the deviation reading.
Encoding moves things slightly. A browser reconstructing a low-bitrate MP3 can shift tracked pitch by a cent or two from the master, which matters when the exactness threshold is six. And measurement is bounded to three minutes from the middle of the file, in a 65 Hz to 1 kHz analysis range, so very high or very low material sits partly outside what is examined.
Using pitch alongside everything else
Pitch is context for a detector score, not a substitute for one. When a scan lands in the ambiguous middle and you want to understand why, cent-exact intonation with zero jitter tells you the pitch was determined by software — which is one of the strongest false-positive clusters in this field. Heavily produced electronic and modern pop material scores higher on average for reasons that have nothing to do with origin, and this is one of the mechanisms behind that.
Read it beside the other free measurements here. Timing shows whether rhythm came from a clock. Structure and repetition show whether passages were copied rather than performed. File metadata carries encoder and editor history. A spectrogram shows band structure and codec history. Stereo field shows mix decisions. No single one is a verdict; together they describe a production history that either matches a stated provenance or does not.
And when the answer genuinely matters — a rights dispute, a contest submission, a label review — the decisive evidence is almost never in the audio. It is the session files, the project history, the timestamps, the stems and the person who can produce them on request. Pitch measurement helps you ask sharper questions of that material. It does not replace it, and any tool claiming otherwise is overselling what arithmetic on a waveform can do.
The short version
A pitch reading tells you what tuning grid a file fits, how closely its notes sit on that grid, and whether pitch moves inside sustained notes. That is genuinely useful: it explains why heavily produced music scores higher on detectors, it exposes speed and shift problems in transfers and re-uploads, and it can test a specific claim that a vocal is unprocessed. What it cannot do is establish origin, because pitch correction is twenty-five-year-old craft and generated vocals routinely imitate human drift. Treat cent-exact intonation as evidence of software in the pitch chain and as a false-positive risk, read it beside timing, structure, metadata, spectrogram and stereo findings, and keep the provenance question with the session files.
Try the free AI music detectorFrequently asked questions
No. Cent-exact intonation means software determined the pitch, which covers pitch correction, software instruments, samplers and generated renders equally. Vocal tuning has been standard commercial practice since the late 1990s, so the measurement identifies processing, not authorship.
More reading
Guide
Tempo, Timing Grids and AI Music: What Quantisation Proves
Quantisation is forty-year-old studio craft. Here is what grid-perfect timing genuinely tells you, and the narrow claims it can settle.
Guide
Stereo Width, Phase and AI Music: What the Image Tells You
Width is the easiest property of an audio file to fake, which makes it the worst to build a provenance claim on.
Guide
Song Structure, Repetition and AI Music: What Loops Prove
Copy-and-paste is a technique, not a confession. Here is what a repetition reading genuinely tells you.