Skip to content

Guide

How to Tell If a Song Is AI Generated: Full Walkthrough

This is the written companion to our long-form walkthrough video: the same ten steps, in the same order, with the readings and the mistakes left in. Follow it end to end and you will reach one of three honest conclusions — likely generated, likely human, or undetermined — and you will know which of your inputs carried the weight.

· 13 min read

Step one: get a copy of the audio worth analysing

Everything downstream is decided here, and almost everybody gets it wrong. The file most people bring to a detector is a screen recording of a phone playing a stream, or a 96 kbps rip, or a fifteen-second clip lifted out of a video. All three destroy exactly the properties analysis depends on: the top octave, the micro-dynamics, and the stereo relationship between channels.

Ask for the original file. If the question matters enough to run a test, it matters enough to request a download rather than a capture. Aim for forty-five seconds or more of continuous music, taken from a section with the full arrangement playing rather than an intro, an outro or a breakdown. If all you can get is a compressed copy, keep going, but hold the result more loosely and say so when you report it.

In the video this step takes twenty seconds and looks like nothing. It is still the step that decides whether the remaining nine are worth doing. A reading taken from a bad copy is not a weak reading; it is a reading of the encoder rather than of the music.

  • Original file over any re-capture, always
  • Forty-five seconds or more, full arrangement, not the intro
  • Note the bitrate and format before you look at any score
  • A poor source does not invalidate the test — it caps how much weight the test can carry

Step two: listen before you measure

Write down what you notice before any tool gives you a number, because a score anchors perception hard. Once you have seen 78% you will start hearing evidence for 78%. Three deliberate passes cost four minutes and make the rest of the process honest.

First pass, vocals only: listen for breath between phrases, consonant grit, and whether repeated lines are phrased identically. Second pass, arrangement: listen for whether instruments change articulation across sections, and whether the transitions between sections land where a player would put them. Third pass, mix: listen for a loud but oddly flat balance, reverb tails that behave the same on everything, and a stereo image that never moves.

None of these is proof, and the honest framing is that listening reliably detects incompetent generated music while performing poorly on competent generated music. That asymmetry is why the listening pass comes second in weight and first in order — it captures your unbiased impression, then gets set aside.

  • Lyrics that scan perfectly and name nothing specific
  • No breath, no consonant noise, identical phrasing on repeats
  • Instruments that never change articulation across a whole track
  • Section joins that are suspiciously exact, or oddly smoothed

Step three: run the scan twice, on different excerpts

One reading is a data point; two readings are information. Run the detector on one excerpt, then again on a different section of the same track, and compare. Our own robustness measurements put the excerpt-to-excerpt shift at around 4.4 percentage points on average, which is the largest source of movement we have measured — larger than re-encoding, and far larger than repeat runs of the same file, which are perfectly deterministic at 0.0 points.

That number is the practical instruction. If two excerpts of the same track land within a few points of one another, the reading is describing the recording. If they disagree by twenty points, the reading is describing whichever section you happened to feed it, and the correct output is undetermined.

The detector runs in the browser, so this costs nothing but a second drag-and-drop. In the video this is the moment the dial appears, and it is also the moment to resist the obvious temptation: the dial is one of five inputs, and it is the fourth most important of them.

Step four: read the reasoning, not the number

Every result carries a breakdown of how it was reached: spectral behaviour, cross-segment tonal agreement, crest factor, spectral-centroid variability, high-band energy ratio and stereo correlation. Reading that panel is what separates a useful conclusion from a superstition, because it tells you which measurement drove the score — and each measurement has a known innocent explanation.

A high reading driven mostly by a hard frequency ceiling usually means you are looking at lossy encoding rather than synthesis. A high reading driven by low crest factor usually means aggressive mastering, which is a production choice made on the majority of commercial releases. A high reading driven by strong section-to-section similarity often means template-driven pop, which is a genre convention, not an origin signal.

So the question to ask of the panel is not 'is the score high' but 'is the driver of this score something only a generator would produce'. In most real cases the answer is no, and that is the finding.

  • Spectral ceiling high → suspect the encoder first
  • Crest factor low → suspect the mastering first
  • Segment agreement high → suspect the genre first
  • Stereo correlation extreme → suspect the mix first

Step five: read the metadata, which is often more revealing

Open the file's tags before you go any further. ID3 and Vorbis fields routinely carry encoder strings, creation timestamps, tool names and comment text that survive an entire distribution pipeline because nobody thought to clear them. A comment field naming a generation tool is worth more than any acoustic reading you will get all day.

The asymmetry matters: metadata traces are suggestive when present and prove nothing when absent, because tags are trivially editable and most upload pipelines strip them anyway. Treat a find as a strong lead and a blank as no information at all.

The exception outranks everything else on this page. Where a file carries Content Credentials or a C2PA provenance manifest, you have cryptographic provenance rather than statistics, and you should stop the acoustic investigation and read the manifest.

Step six: ask for provenance, which decides most real cases

This is the step people skip because it involves talking to a person. It is also the step that resolves cases. Ask for stems, a project file with a plausible edit history, an unmastered rough, dated drafts, or a phone video from the tracking session. Any one of those outweighs everything you measured.

Frame it as a request, not an accusation — and remember that AI-assisted music is legal, widely used and often disclosed. Plenty of artists will simply tell you what they used, at which point the investigation is over and nobody has been insulted. A working musician can produce stems in minutes; a volume uploader running a generation pipeline usually cannot produce anything at all, and that gap is the actual signal.

If the answer is a refusal with no explanation, you still do not have proof of generation. You have an absence of documentation, which is a legitimate basis for a decision in a commercial context and not a basis for a public accusation.

  • Cryptographic provenance or a direct admission — decisive
  • Stems, project files, dated drafts, session video — very strong
  • A cooperative, specific answer — usually the end of the matter
  • A refusal — an absence of evidence, not evidence of generation

Step seven: check the release pattern

Context is the second-strongest input and takes about three minutes. Look at the artist's catalogue rather than the single track: how much material, released how quickly, across how many unrelated genres, with what production consistency, on an account how old, with what live footage attached.

Twelve albums in a month across drill, ambient and country, all with identical mastering and no live performance anywhere, tells you more than any spectrum. Equally, a decade of releases, gig photos and a coherent stylistic thread makes a high acoustic reading much likelier to be a production artefact than an origin signal.

Nothing here is conclusive either. Prolific human artists exist, and generated catalogues can be built slowly on purpose. But context is cheap, hard to fake at scale, and it is what turns a lone percentage into an actual case.

Step eight: weigh the five inputs in the right order

Now put the pieces in order of evidential weight: provenance first, context second, metadata third, acoustic analysis fourth, subjective listening fifth. That ordering is deliberately unflattering to the tool on this site, and it is the honest one. Acoustic analysis is the input you can obtain in thirty seconds, which is exactly why it gets over-weighted by everyone who reaches for it.

A useful discipline: write your conclusion as a sentence naming its strongest input. 'Likely generated, based on a refusal to supply stems, forty releases in six weeks, and two consistent high readings' is a defensible finding. 'Likely generated, 81%' is not a finding at all — it is a screenshot.

Where the inputs point in different directions, the answer is undetermined, and undetermined is a real result rather than a failure. Most careful investigations of a single track end there, and saying so protects both you and the artist.

  • High reading plus solid provenance → the detector is reading production style
  • Low reading plus no provenance at all → says very little; post-production removes artefacts
  • Consistent readings plus anomalous release pattern → the strongest acoustic-led case available
  • Inputs in conflict → undetermined, stated plainly

Step nine: know the two ways you will be wrong

False positives cluster in a predictable place: electronic and synthesiser-based human music. In our own measurements of human synthesiser renders, three of twelve read above 70% — not because the tool malfunctioned, but because a fully programmed, quantised, loudness-maximised human production genuinely shares measurable properties with generated audio. If the track is techno, trance, hyperpop or a game soundtrack, raise your threshold before you conclude anything.

False negatives cluster just as predictably: anything that has been through human post-production. Re-singing a topline, re-amping a guitar, running the mix through analogue gear, or simply having a competent mastering engineer touch it removes most of what remains detectable. A confident low reading on a polished hybrid is the failure mode nobody notices, because it agrees with what people want to hear.

Both of these are properties of the field rather than of one product. Any tool that hides them from you is selling confidence instead of information, which is the single most useful thing to know when comparing detectors.

Step ten: report it in a way that survives scrutiny

State the conclusion, the strongest input, the readings with their excerpts, the source quality, and the limits — in that order, in about five lines. If you are exporting the PDF report, add the two things it cannot know: where the file came from, and what the artist said when asked.

Keep the language proportionate to the evidence. 'The audio shows characteristics associated with generated music, and no provenance was supplied on request' is accurate. 'This song is AI' is not, and in a dispute it is the sentence that will be quoted back at you.

Then stop. A detector output is a reason to ask a question, not an answer to one — and in a moderation queue, a classroom, a label inbox or a comments section, the defensible action is always to request documentation rather than to publish a percentage.

  • Conclusion, strongest input, readings, source quality, limits
  • Name the excerpts you tested and their agreement
  • Never attribute output to a named generator from audio alone
  • Ask for project files before you say anything publicly

The short version

Ten steps, five inputs, one honest ordering: provenance, context, metadata, acoustic analysis, then your own ears. Get a real copy of the file, listen before you measure, scan two excerpts and read the reasoning behind the number, ask for project files, and be willing to end at undetermined — which is where most single-track investigations honestly end.

Try the free AI music detector

Frequently asked questions

  • About fifteen minutes if you have the file and the artist answers, of which the detector accounts for under a minute. The listening pass and the release-pattern check take the most time and carry more weight than the score.

More reading