Skip to content

Encyclopedia

Forensic audio encyclopedia

Every measurement this tool makes, defined in plain language — with an explicit note on what else produces the same reading in genuine human recordings. If a term is listed as a weak signal, that is our assessment of it, not a disclaimer we hid at the bottom.

Spectral

Spectral centroid

Moderate signal

The centre of mass of a sound's frequency spectrum — a numerical measure of brightness.

Spectral centroid is the amplitude-weighted average frequency present in a short window of audio. A track with lots of cymbal and sibilance has a high centroid; a bass-heavy dub mix has a low one. On its own the absolute value says almost nothing useful about origin, because it mostly describes genre and arrangement.

What is informative is how the centroid moves. In a performed recording it jitters constantly: a snare hit pushes it up for 40 milliseconds, a sustained pad pulls it down, a consonant spikes it. The variance of that movement over time is a rough proxy for how much genuine acoustic event variety a recording contains.

Generated and heavily processed audio often shows lower centroid variance across a track than a comparable live performance. It is a difference of degree, not of kind, which is exactly why it belongs in a weighted ensemble rather than in a rule.

How this tool uses it: Measured per analysis window; the variance across windows contributes to the spectral component of the score at moderate weight.

What else produces the same reading

  • Multiband compression and limiting flatten brightness movement in human recordings.
  • Programmed electronic music has genuinely low event variety.
  • Lossy encoding removes high-band content and shifts the centroid downwards.

Spectral flatness

Moderate signal

How noise-like versus tone-like a spectrum is, expressed as a single ratio.

Spectral flatness compares the geometric mean of a spectrum to its arithmetic mean. A pure sine wave has flatness near zero; white noise has flatness near one. Real music sits between the two and moves around as the arrangement changes.

It is useful in detection mostly as a texture descriptor for the upper bands. Synthesised high-frequency content can be either too smooth — a modelled hiss with no structure — or too structured, showing regular energy where a real cymbal would show chaos.

How this tool uses it: Computed on the upper spectrum and combined with high-band energy distribution; contributes at moderate weight and is discounted on low-bitrate files.

What else produces the same reading

  • Noise reduction and de-noising plug-ins smooth the same bands.
  • Vinyl rips and tape sources add broadband noise that raises flatness.
  • Any codec below roughly 192 kbps rewrites this measurement entirely.

Spectral ceiling

Moderate signal

The frequency above which a recording contains essentially no energy.

Many generative models are trained on band-limited representations, so their output simply stops at a particular frequency, often with a sharp shelf rather than the gradual roll-off of a microphone recording. Seen on a spectrogram it is a horizontal line with nothing above it.

This used to be one of the most cited AI-audio giveaways. It is far less useful than its reputation suggests, because every lossy codec produces the same shape. A 128 kbps MP3 of a live orchestra also has a hard ceiling near 16 kHz.

The honest position: a hard ceiling on a verifiably lossless file is meaningful evidence. A hard ceiling on a file of unknown provenance is close to worthless.

How this tool uses it: Detected and reported, but automatically down-weighted whenever other evidence indicates lossy encoding. It can never carry a result alone.

What else produces the same reading

  • MP3, AAC and Opus encoding at any moderate bitrate.
  • Old or low-quality source recordings.
  • Deliberate low-pass filtering as a production choice.

Dynamics

Crest factor

Moderate signal

The ratio between a signal's peak level and its average level.

Crest factor quantifies how much headroom sits above the average loudness. A sparse acoustic recording has a high crest factor; a brick-walled modern master has a low one. It is one of the cheapest measurements available and one of the most misused.

Generated tracks frequently arrive already loudness-maximised, so they cluster at low crest factors. So does most commercial music released since about 2005. The measurement separates 'has been maximised' from 'has not been maximised' — it does not separate machine from human.

How this tool uses it: Contributes to the dynamics component. Its weight is automatically reduced when spectral evidence disagrees with it, because reference-conditioned generators can inherit human dynamics.

What else produces the same reading

  • Loudness-war mastering on human records.
  • Streaming normalisation applied after the fact.
  • Reference-conditioned generation copying a human envelope.

Dynamic range

Moderate signal

The spread between the quietest and loudest meaningful passages of a recording.

Where crest factor is a short-window measurement, dynamic range describes the whole piece: does the second chorus actually hit harder than the first verse, or has everything been levelled to the same loudness?

Automatic generation with automatic mastering tends to produce arrangements that are dynamically flat at the macro scale. Human records vary enormously by genre — a jazz trio and a metal record are not comparable — so the measurement is only interpretable alongside the rest of the evidence.

How this tool uses it: Measured across the full file and across segments; feeds both the score and the confidence estimate.

What else produces the same reading

  • Genre. Electronic and pop production is legitimately flat.
  • Streaming loudness normalisation.
  • Short clips, which may not contain a dynamic contrast at all.

Stereo

Stereo field behaviour

Moderate signal

How energy is distributed and moves between the left, right and side channels.

Average stereo width is nearly useless for detection — it is a mixing decision. What carries information is how the side channel behaves over time relative to the mid channel.

A recording made with real microphones in a real space has correlated, physically plausible movement: a room reflects, instruments occupy positions, reverb decays into the sides. Synthesised stereo can be plausible too, but it is often produced by a widening process applied uniformly, which shows up as side-channel energy that tracks the mid channel too closely or too evenly.

How this tool uses it: Mid/side energy ratios are computed per window; the variability of that ratio across windows is the feature, not the average.

What else produces the same reading

  • Stereo wideners and mid/side processing on human mixes.
  • Mono or near-mono sources of any origin.
  • Upmixed mono, which produces artificial but human-authored stereo.

Phase coherence

Weak signal

Whether the two channels of a stereo file relate to each other in a physically plausible way.

Phase relationships encode a great deal about how audio was captured — a pair of microphones at a distance produces predictable inter-channel delays. Reconstructed or synthesised audio does not necessarily obey those constraints.

In practice this is a weak signal for detection, because the vast majority of released music is not a raw microphone capture. Multitrack production, virtual instruments and stereo effects break the physical assumptions long before any AI is involved.

How this tool uses it: Observed only indirectly through mid/side behaviour. No standalone phase claim appears in the report.

What else produces the same reading

  • Virtually all modern multitrack production.
  • Software instruments and synthesised sources.
  • Reverb and delay plug-ins.

Temporal

Segment agreement

Strong signal

How consistently independent sections of a track produce the same measurement.

The engine splits a file into overlapping windows and analyses each one on its own before combining the results. Agreement between those windows is arguably the most valuable single piece of information the analysis produces — and it feeds confidence rather than probability.

If ten segments all point the same way, the reading is stable and the confidence level reflects that. If four say one thing and six say another, the honest output is 'inconclusive', and that is what the report shows. A tool that averages disagreement into a confident-looking number is hiding its own uncertainty.

How this tool uses it: Primary input to the confidence level. Sufficiently low agreement forces an inconclusive outcome regardless of the mean probability.

What else produces the same reading

  • Very short clips contain too few segments to measure agreement usefully.
  • Tracks with dramatic arrangement changes legitimately disagree with themselves.
  • Mashups and remixes combine sources of different origin.

Transient density

Moderate signal

How many sharp attack events occur per unit of time, and how varied they are.

Percussive attacks, pick noise and consonants are transients: short bursts of broadband energy. Human performance produces transients that vary in amplitude and micro-timing on every repetition, because human motor control is not a sequencer.

Quantised and generated material tends to produce transients that repeat with unusually tight timing and amplitude. As a detection feature this overlaps almost completely with 'was this programmed rather than played', which is a different question from 'was this generated'.

How this tool uses it: Contributes to the temporal component at moderate weight, alongside brightness variability.

What else produces the same reading

  • Quantised programming in every electronic genre.
  • Drum replacement and sample triggering in human production.
  • Heavy compression, which levels transient amplitude differences.

Micro-timing variation

Weak signal

The tiny deviations from a strict grid that characterise human performance.

A drummer pushing the backbeat by six milliseconds is expressive; a sequencer placing it exactly on the grid is not. Measuring this reliably from a mixed stereo file, without stems and without a known tempo map, is genuinely difficult, and claims about it should be treated with suspicion.

We list it here because it is frequently cited in popular explanations of AI music detection as though it were easy to measure. It is not, and this engine does not claim to measure it directly.

How this tool uses it: Not used as a standalone feature. Only its coarse consequences are visible, through transient density and segment agreement.

What else produces the same reading

  • Quantised human productions have no micro-timing variation either.
  • Tempo drift in live recordings complicates any grid-based measurement.

Provenance

Codec artefacts

Moderate signal

Traces left by lossy encoders such as MP3, AAC and Opus.

Lossy codecs discard information the encoder's psychoacoustic model believes is inaudible. What remains carries traces: a bandwidth ceiling, quantisation patterns in the frequency domain, pre-echo around sharp transients, and joint-stereo behaviour in the upper bands.

For detection purposes codec artefacts are mostly a nuisance, because they overwrite the fine detail several features depend on. But identifying them is valuable in itself: knowing a file has been re-encoded tells you how much to trust the rest of the reading.

How this tool uses it: Encoding evidence is used to modulate the weight of spectral features and to lower confidence, not to raise or lower the probability directly.

What else produces the same reading

  • Multiple generations of re-encoding compound each other.
  • Transcoding a lossy file to WAV hides the file extension but not the artefacts.

Metadata and provenance

Strong signal

Non-acoustic evidence: file tags, upload history, project files, platform records.

The strongest available evidence about a track's origin is almost never acoustic. Project files with edit history, session stems, dated platform uploads, distributor records and a creator who can explain their own arrangement outweigh any spectral measurement.

Metadata is also trivially removable and forgeable, so its absence proves nothing while its presence proves quite a lot. Any serious verification process should start here and use acoustic analysis as corroboration.

How this tool uses it: Deliberately not used. Classification is performed on audio content only, and our process never reads or stores file metadata.

What else produces the same reading

  • Metadata is stripped by most upload pipelines automatically.
  • Tags can be written by anyone with a free tool.

Audio watermarking

Strong signal

A deliberate, usually inaudible marker embedded by the generator at synthesis time.

Some platforms embed watermarks so their own output can be recognised later. When a watermark is present and the reading key is available, this is far more reliable than any statistical inference — it is a designed signal rather than an inferred one.

The catch is that watermarks are proprietary, detectable only by the party that issued them, and often survive poorly through re-encoding, pitch shifting or re-recording. From an independent tool's position they are unreadable.

How this tool uses it: Not read. This tool performs acoustic analysis only and makes no watermark claims in either direction.

What else produces the same reading

  • Absence of a detectable watermark says nothing, since you cannot read most of them anyway.
  • Re-encoding and re-recording can destroy watermarks.

Method

FFT and windowing

Moderate signal

The transform that turns a slice of audio into a frequency spectrum, and the taper applied first.

Every spectral measurement on this site starts with a Fast Fourier Transform of a short slice of audio. Cutting a slice out of a continuous signal creates artificial discontinuities at the edges, which smear energy across the spectrum — so each slice is multiplied by a window function, here a Hann window, that tapers it to zero at both ends.

The choice of window length is a direct trade-off: longer windows give finer frequency resolution and coarser timing, shorter windows the reverse. This is not a detail — it determines which artefacts a detector can see at all.

How this tool uses it: Hann-windowed FFT with overlapping frames is the foundation of every spectral feature the engine computes.

What else produces the same reading

  • Not applicable — this is a method, not a signal.

Ensemble scoring

Strong signal

Combining several independent detectors into one probability with a confidence level.

No single feature separates generated from human music. An ensemble runs several independent analyses over the same audio and combines their outputs with configured weights, so that a strong reading from one model cannot on its own produce a confident verdict.

Crucially, the ensemble tracks disagreement. When components diverge, the confidence level falls and the outcome can be reported as inconclusive. That is a feature: the alternative is a system that always sounds sure.

How this tool uses it: Weighted combination over the configured components, with per-component weights held in a configuration file rather than hard-coded in the analysis path.

What else produces the same reading

  • Correlated components can look like agreement while sharing the same blind spot.
  • Weights are judgement calls until a published benchmark constrains them.

Probability versus confidence

Strong signal

Two different numbers: how likely, and how much the evidence can be trusted.

Probability answers 'how machine-like do these measurements look?'. Confidence answers 'how much should you rely on that estimate for this particular file?'. They move independently, and conflating them is the most common way detection tools mislead people.

A 40-second lossless upload with strong segment agreement can produce a 62% probability at high confidence. A 12-second 96 kbps clip can produce 84% at low confidence. The second number looks more alarming and means less.

How this tool uses it: Both are always shown. Where confidence is too low to support any claim, the outcome is labelled inconclusive rather than being reported as a score.

What else produces the same reading

  • Short clips inflate apparent certainty while removing the evidence behind it.

Related