Skip to content

Guide

How to Read a Spectrogram: AI Music Clues and Codec Myths

A spectrogram is the most honest picture you can take of a recording, and the most widely misread. It shows exactly where energy sits in frequency and time, which makes codec history obvious and generative origin far less obvious than the internet suggests. This is how to read one properly: what the shapes mean, which ones carry information about AI, and which famous 'tells' are nothing more than an MP3 doing its job.

· 12 min read

What a spectrogram actually shows

A waveform plots amplitude against time: loud and quiet, nothing else. A spectrogram adds the dimension that matters for forensic work. Time runs left to right, frequency runs bottom to top, and brightness encodes how much energy sits at that frequency in that instant. Everything you can see in a spectrogram is already in the audio — the transform reveals nothing new, it simply reorganises the signal into a layout that human eyes are good at scanning.

The underlying operation is a short-time Fourier transform. The file is cut into overlapping frames a few tens of milliseconds long, each frame is tapered by a window function so its edges do not create artificial clicks, and an FFT converts each frame into a column of frequency bins. Stack the columns and you have the image. Our viewer analyses up to thirty seconds taken from the middle of a track, because intros and fade-outs are the least representative part of any recording.

Two consequences follow from that construction and they shape every reading you will do. First, there is a resolution trade-off: short frames give you sharp timing and blurry frequency, long frames give you sharp frequency and smeared timing. You cannot have both, so a transient may look like a vertical smudge rather than a line. Second, the picture is logarithmic in perception but often linear in layout, which means the musically busy bottom two octaves get squeezed into a thin band at the base while the sparse top end takes up half the height.

The vocabulary of shapes

Before you can look for anything meaningful you need to be able to name what is on screen. Almost every spectrogram is built from six recurring shapes, and once you can identify them the picture stops being noise and starts being a description of the arrangement.

  • Horizontal lines: sustained pitched content. A bass note, an organ pad or a held vowel appears as a stack of evenly spaced lines — the fundamental plus its harmonics.
  • Vertical lines: transients. Kick drums, snare hits, guitar picks and consonants dump energy across many frequencies in a very short window.
  • Wavy or drifting lines: vibrato, pitch bends, tape flutter, slides. Any deviation from a machine-steady pitch shows up here, which is why this shape matters later.
  • Broad clouds: noise-like content. Cymbals, breath, room reverb, distortion and hiss occupy wide bands rather than discrete lines.
  • Sharp horizontal edges: filtering. A clean boundary above which everything goes dark is a lowpass filter or a codec cut-off, not a musical event.
  • Repeating rectangles: loops. If a two-bar block of the image is pixel-similar to the block before it, the same audio was pasted rather than performed again.

The band edge, and the myth built on it

The single most visible feature in most spectrograms is the point where the image goes black at the top. In lossy files this is a hard, ruler-straight line, and its height is set almost entirely by the codec and bitrate rather than by anything musical. A 128 kbps MP3 typically stops somewhere between 15 and 16 kHz, a 192 kbps file a little higher, a 320 kbps file around 20 kHz, and AAC and Opus place their own characteristic ceilings. Uncompressed WAV and FLAC have no ceiling at all beyond the Nyquist limit of half the sample rate.

This is the feature that launched a thousand confident misreadings. 'The spectrum stops at 16 kHz, therefore it is AI' is a claim you will see repeated in comment sections and it is wrong in the plainest possible way: it is a claim about the file's compression history, not about how the music was made. Send any human recording through a streaming platform, a messaging app or a screen capture and it acquires exactly the same shelf.

Early generative systems did produce narrow output, because they synthesised into compressed or bandwidth-limited representations, and that history is why the myth has such staying power. Current systems mostly do not. Meanwhile the world's supply of ordinary human music has been re-encoded so many times that a limited band edge is now the default state of audio in circulation. A feature shared by nearly everything cannot discriminate between anything.

What the band edge is genuinely good for is telling you how much you should trust everything else. If the file stops at 15 kHz, then any measurement of high-frequency behaviour is measuring the codec, and the honest response is to weight it near zero. That is precisely how our detector treats it, and it is why the viewer reports the visible edge and the share of energy above 12 kHz side by side.

What a spectrogram can genuinely contribute

Set aside origin for a moment. A spectrogram is excellent at answering questions about construction, and construction questions are often the ones that actually resolve a dispute.

Repetition is the strongest example. Human performances of the same section differ every time — a snare lands two milliseconds early, a vocal breath sits in a different place, the room decays slightly differently. Copy-paste arrangement, whether done by a producer in a DAW or by a model that reuses its own output, produces blocks that are visually identical rather than merely similar. You can see that at a glance in a way you cannot hear reliably.

Editing history is the second. Splices show as discontinuities running the full height of the image, where reverb tails are cut off mid-decay. Pitch correction sometimes appears as suspiciously flat horizontal lines where vibrato has been ironed out, or as tiny stair-steps at note boundaries. Noise reduction leaves a characteristic gating pattern in the quiet gaps. None of these prove AI, but all of them tell you the recording has been processed, which is exactly what you need to know before reading a probability.

Third, a spectrogram lets you see the mix in a way that explains a score. If everything in the picture is bright and dense with no dark gaps anywhere, you are looking at heavy limiting — and that tells you why the crest-factor measurement came back low, without you having to attribute it to generation.

  • Blocks that repeat pixel-for-pixel: reused audio rather than a re-performance
  • Full-height discontinuities with truncated reverb tails: edit points
  • Vibrato lines that flatten out mid-phrase: pitch correction
  • Gating patterns in the silences: noise reduction or a gate
  • A picture with no dark space at all: severe limiting, which suppresses dynamic-range measurements
  • A hard shelf plus near-zero energy above it: lossy encoding, so discount every high-band reading

Where visual inspection stops working

It would be convenient if generative pipelines left a visible signature. Some once did. Certain vocoder and diffusion decoders produced regular horizontal striping in the upper mid range, or a faint checkerboard texture in quiet passages, and for a while those were reasonable things to look for. Model generations turn over in months, and each turnover has moved output further away from anything you can name on sight.

The deeper problem is that mastering and generation do overlapping things to a spectrogram. Both flatten dynamics. Both produce dense, wide, consistently loud pictures. Both can narrow the top end. A heavily limited electronic master and a generated track can look close to identical, which is the same confound that makes acoustic scoring hard rather than a separate weakness of the visual method.

There is also a plain statistical point that gets lost in enthusiasm for the image. Looking at a spectrogram without a reference distribution means you have no idea whether what you are seeing is unusual. 'This looks strange to me' is not a measurement. Our detector exists because turning these behaviours into numbers, comparing them against how real material behaves, and reporting a range with a confidence level is the only way to say anything defensible.

So use the viewer as an explanation tool and a construction check, not as an origin oracle. When it shows you a 15 kHz shelf, you have learned why the high-band signal was discounted. When it shows you four identical bars, you have learned something real about the arrangement. When it merely looks odd, you have learned that spectrograms of music often look odd.

A practical reading order

The way to avoid fooling yourself is to look at the same things in the same order every time, and to write down what you saw before you look at any score. Anchoring is the dominant failure mode in this work: if you already believe a track is generated, ambiguous shapes will look like confirmation.

  • Start at the top and find the band edge. Note its frequency and whether the cut is hard or gradual. That number sets how much of the rest of the image you can trust.
  • Scan the bottom two octaves for structure — is the low end a series of clean harmonic stacks or a continuous smear?
  • Look for repeated blocks across the whole width. Compare bar to bar rather than second to second.
  • Check the vertical lines. Do transients vary in strength and spacing, or do they arrive with metronomic identical intensity?
  • Check pitched lines for movement. Total steadiness across a long held note is worth noting.
  • Look at the dark space. How much of the picture is quiet? None at all means heavy limiting.
  • Only now run the detector, and read the measurements against what you already wrote down.

Why this runs in your browser

The viewer decodes the file with the Web Audio API, computes the transform in the page, and paints it to a canvas. The audio never leaves your device: nothing is uploaded, nothing is stored, and closing the tab destroys everything. That is the same architecture as the detector and the metadata checker, and it is not incidental.

Unreleased material is the normal case for the people who need these tools. A label doing intake, an artist checking their own mix, a moderator handling a submission — none of them should have to post a private file to a server to find out where its spectrum stops. Local processing means the honest answer to 'what do you do with my audio' is that there is nothing to do with it.

The trade-off is real and worth stating. Browser processing means larger files take longer, very long tracks are sampled rather than analysed end to end, and the picture is capped at what a canvas can usefully show. We think a slightly coarser image you can trust with an unreleased master beats a finer one that requires an upload.

The short version

A spectrogram tells you what happened to a file, not who or what made the music. Read the band edge first and treat it as a note about compression history that tells you how much the high-frequency evidence is worth. Then use the picture for what it is genuinely strong at: repeated blocks, edit points, flattened vibrato and the absence of dark space. Save the origin question for measurements with a reference distribution behind them, and be suspicious of anyone who claims to spot AI by eye.

Try the free AI music detector

Frequently asked questions

  • Not reliably, and not from the features most often cited. Current generative output does not carry a visible signature that survives ordinary mastering, and the classic 'hard cut-off at 16 kHz' tell is a property of lossy compression that human recordings acquire routinely. A spectrogram is genuinely useful for spotting repetition, edit points and heavy limiting, all of which describe how a recording was constructed rather than whether a model made it.

More reading