Skip to content

Guide

LUFS, Loudness and AI Music: What a Meter Really Tells You

Crushed dynamics get cited as an AI tell more often than almost anything else, and it is one of the weakest claims in the whole field. Loudness is a mastering decision, made after the music exists, by a human or a preset or a generator's own output chain. This is what LUFS actually measures, how streaming normalisation changes what listeners hear, and why a loudness reading belongs in a mastering conversation rather than a provenance one.

· 12 min read

What LUFS measures, precisely

LUFS stands for Loudness Units relative to Full Scale, and it is defined by ITU-R BS.1770 with the gating rules that EBU R 128 layers on top. The measurement starts by filtering the audio through a two-stage 'K-weighting' curve: a high shelf of roughly +4 dB above about 1.7 kHz, and a high-pass near 38 Hz. Those two filters are a crude model of how sensitive human hearing is across the spectrum, and they are the reason a bass-heavy track and a bright track with identical peak levels read differently.

After filtering, the signal is chopped into 400 ms blocks with 75% overlap and each block's mean square is computed, channel-weighted, and converted to a loudness value. Then two gates run. The absolute gate throws away everything below −70 LUFS, which removes silence and room tone. The relative gate computes the mean of what survives and discards anything more than 10 LU below it, which removes quiet passages that would otherwise drag the average down. What remains, averaged, is the integrated loudness — one number for the whole programme.

Two other numbers matter for reading a master. Loudness range, or LRA, comes from the distribution of 3-second short-term values between roughly the 10th and 95th percentiles, and it describes how much the level moves between sections. Crest factor is simply peak minus RMS in decibels, and it describes how far transients stick out above the average. Neither is a quality judgement; both are descriptions of a shape.

Why platform normalisation makes loudness wars pointless

Every major streaming service now normalises playback. Spotify, YouTube Music, Amazon and Tidal aim for around −14 LUFS; Apple Music's Sound Check sits nearer −16; broadcast delivery under EBU R 128 is −23. The practical effect is that if you master at −6 LUFS, the platform turns you down by roughly 8 dB before anyone hears it. You do not arrive louder. You arrive at the same level as everyone else, with less dynamic range left to show for it.

There is an asymmetry worth knowing: some services turn loud masters down but will not turn quiet ones up. YouTube behaves this way. So a very quiet master can genuinely sound weaker there than the tracks around it, while a very loud one gains nothing anywhere. That is the entire argument for landing somewhere in the −9 to −14 window for most contemporary music, and for leaving about a decibel of true-peak headroom so lossy encoding does not push inter-sample peaks past full scale.

None of this has anything to do with how the music was made. It is a delivery question that applies identically to a string quartet recorded in a hall and a fully generated track exported from a browser.

The 'AI music is over-compressed' argument, examined

The claim usually runs: generated tracks come out loud and flat, therefore a loud flat master is evidence of generation. Both halves are shaky.

Generated output is often loud and flat because most generators apply a limiter on export and because their training data is modern commercial music, which is itself loud and flat. But so is essentially every contemporary pop, hip-hop, EDM and metal release from the last twenty years, mastered by humans who were competing for perceived loudness before normalisation made that pointless. A −7 LUFS master with a 6 dB crest factor describes a huge portion of the charts. As a discriminator, it is close to useless.

The converse fails too. Anyone can drop a generated track into a DAW and re-master it to −16 LUFS with 12 dB of crest factor in about four minutes. Loudness is the single easiest property of a file to change, which makes it the single worst property to build a provenance claim on. Any signal that a two-step export can erase was never carrying much information.

Where loudness genuinely helps is in the opposite direction: as context for reading a detector score. A heavily limited master has had transient detail flattened, and transient behaviour is one of the things detection methods look at. Knowing a file reads −6 LUFS with a 5 dB crest factor tells you why a score might be pushed upward for reasons that have nothing to do with origin — a false-positive risk factor, not evidence.

Reading your own measurements

Some rough bands for interpreting what our loudness tool reports, with the caveat that genre changes what is normal.

  • Integrated above −8 LUFS: very loud. Every platform will attenuate it. Usually a sign the limiter is doing structural work rather than catching peaks.
  • −8 to −11 LUFS: loud, and typical of modern commercial pop and electronic releases. Harmless, but the loudness ends at the platform.
  • −11 to −17 LUFS: the window most services aim for, so playback level changes little. The safest place for a master to sit.
  • Below −17 LUFS: quiet. Spotify and Apple will turn it up; YouTube will not, so it can sound thin in context there.
  • Crest factor under 8 dB: heavily limited. Common, and a reason to be sceptical of a high detector score on the same file.
  • Crest factor above 12 dB: dynamic. Typical of acoustic, jazz, classical and lightly mastered material.
  • LRA under 3 LU: level barely moves between sections. Can be an arrangement choice, a mastering choice, or just a short excerpt.
  • Sample peak at or above −0.05 dBFS: no headroom. Lossy encoding can push true peaks over 0 dBFS and clip on some players.

Sample peak, true peak, and why the difference matters

A sample-peak meter reports the largest sample value in the file. A true-peak meter oversamples first — usually 4× — to estimate where the reconstructed analogue waveform actually goes between samples. Those inter-sample peaks can sit a few tenths of a decibel, occasionally more than a decibel, above the highest sample.

This is why delivery specs ask for −1 dBTP rather than −0.1 dBFS. An MP3 or AAC encoder reshapes the waveform; a file that reads exactly 0.00 dBFS on samples can decode to something that overshoots and clips in a consumer DAC. Our browser-side meter measures sample peak, so treat its reading as a floor and assume the true peak is slightly higher.

Why this runs in your browser

The loudness tool decodes your file with the browser's own audio decoder and runs the BS.1770 filters in JavaScript on your machine. Nothing is uploaded, nothing is stored, and closing the tab is the whole deletion process. For unreleased masters and material under NDA that is not a nicety — it is the only version of the tool that is usable at all.

The trade-offs are honest ones. Very long files are measured from a bounded excerpt rather than end to end, so a two-hour DJ set will report the excerpt's loudness rather than the programme's. Peak measurement is sample-based, not oversampled. And the browser's decoder is the one deciding how your lossy file becomes samples, which means readings can differ by a few tenths from a desktop meter using a different decoder. For mastering decisions and for sanity-checking a detector score, that resolution is more than enough; for contractual broadcast delivery, use a certified meter.

Where a loudness check fits in a provenance workflow

Treat it as a context step, never as a verdict step. Run the detector first and read the range and confidence. If the estimate is high, check loudness and crest factor before you act: a heavily limited master is a known source of upward pressure, and finding one should widen your uncertainty rather than confirm your suspicion.

Then keep going with the checks that actually carry information about construction: file metadata for encoder and software fields, a spectrogram for repeated blocks and edit points, a stability check across excerpts to see whether the track is homogeneous, and — when the stakes are real — a direct request for stems, session files or a writer credit. Loudness earns its place in that sequence by explaining anomalies, not by producing them.

The short version

LUFS describes a mastering decision, not an origin. Platform normalisation has already removed the reward for loud masters, so aim for the −9 to −14 window with about a decibel of true-peak headroom and stop competing. When you are trying to work out whether a track was generated, use loudness and crest factor to explain why a detector score might be inflated — a crushed master flattens the transient detail detection relies on — and leave the actual provenance work to metadata, spectrograms, stability across excerpts, and asking the artist.

Try the free AI music detector

Frequently asked questions

  • No. Loudness and dynamic range are set after the music exists, by a mastering chain that can be human, a preset or a generator's export limiter. Most commercial releases of the last two decades are loud and flat, and any generated track can be re-mastered to be dynamic in minutes. It is a mastering fact, not a provenance signal.

More reading