Skip to content

Guide

Stereo Width, Phase and AI Music: What the Image Tells You

"It sounds too wide, too glossy, too evenly spread — must be AI." Stereo width is one of the most confidently misread properties in this whole argument, and one of the cheapest to fake. This is what correlation, mid/side balance and mono compatibility actually measure, what they tell you about a mix, and the narrow, honest role they play when you are trying to work out where a track came from.

· 12 min read

Mid, side, and what "width" really is

A stereo file is two channels, but it is more useful to think of it as two derived signals. The mid signal is what the channels share — the sum, scaled — and the side signal is what they differ by, the difference. Anything panned dead centre lives entirely in the mid. Anything that appears in only one channel, or appears in both with different timing or tone, contributes to the side.

Width, in the measurable sense, is the share of total energy carried by that side signal. It is not a property of the music. It is a property of how the music was placed and processed: panning decisions, room mics, delayed doubles, reverb returns, and whatever widening plugin sat on the master bus. Two mixes of the same performance can differ by a factor of five in side energy without a single note changing.

Most commercial masters land somewhere between roughly 2% and 15% side energy. Below that you have a near-mono image, which is normal for older recordings, spoken material and some deliberately tight productions. Well above it and you are usually looking at a widener, a Haas-effect delay, or a mid/side balance that has been pushed past what the mix can support.

Inter-channel correlation and what its sign means

Correlation compares the two channels sample by sample and returns a number between −1 and +1. It answers a single question: to what extent is one channel a copy of the other?

Near +1 means they are almost the same signal. The image is centred, and summing to mono changes essentially nothing. Around +0.5 to +0.8 is where a great deal of well-behaved pop and rock mastering sits: a solid centre with real width around it. Near 0 means the channels carry largely independent content, which is entirely normal for wide ambient, orchestral and electronic material recorded or synthesised with genuinely different signals on each side.

Negative values are the interesting case. They mean that when the channels are summed, parts of the signal cancel. Occasionally that is deliberate — some widening processes work by injecting anti-correlated content. Far more often it means a polarity inversion crept in somewhere: a mis-wired cable, a flipped plugin switch, a sample imported with inverted phase, a bass double that fights its own source. A negative overall correlation on a finished master is almost always a fault, not a style.

Mono compatibility, and why it still matters

A lot of real listening sums to mono or something close to it. Single-speaker phones, most smart speakers, many Bluetooth speakers, club systems with mono-summed subs, some broadcast chains, and shopping-centre PA. If your master loses several decibels when the channels are summed, everyone on those systems hears a thinner and quieter version of the mix you signed off.

The mono-fold figure is exactly that: the level change when left and right are summed. Zero means nothing cancels. A drop of a decibel or so is unremarkable. Three or more decibels means a meaningful part of the mix is disappearing, and it is worth finding out which part.

The usual culprit is low frequency. Wide bass is the single most common cause of mono cancellation, which is why most engineers keep everything below roughly 150 Hz near the centre — not out of superstition, but because a kick and a bass that cancel take the foundation of the track with them. This is what the per-band table on the tool is for: it shows how much side energy each band carries, so you can see whether your width lives in the cymbals and reverb where it is harmless, or in the low end where it is not.

The "too wide, therefore AI" claim

There is a grain of truth underneath it. Some generated output does have a characteristically even stereo picture: a wide, homogeneous spread that barely changes between sections, because the track was rendered in one pass as a finished stereo mix rather than assembled from separately recorded and separately panned parts. A human arrangement usually shows movement — a verse that narrows, a chorus that opens out, a bridge with a different spatial signature — because someone made those decisions one at a time.

But that description also fits an enormous amount of ordinary human production. Anything mixed quickly with a stereo widener across the master, anything built from wide stock loops, anything mastered by an automated service: same flat, glossy, unchanging image. And the reverse is trivially available. Re-mastering a generated track to have narrow verses and wide choruses takes minutes, and simply running any file through a mid/side adjustment changes every number on this page while leaving the audio's actual origin untouched.

That is the fatal problem with width as evidence. It is one of the easiest properties of an audio file to alter, and a signal that a two-minute edit can erase or manufacture was never carrying much information about how the music was made.

The role these numbers do play

Read them as context for a detector score rather than as an input to it. Detection methods read spectral and transient detail, and the stereo image changes what that detail looks like. Both extremes matter, in opposite ways.

A collapsed, near-mono image removes the inter-channel differences that some analysis paths rely on, so there is simply less to work from — and less material generally means a wider, less confident range. An aggressively widened image is the more dangerous case: the processing that creates artificial width introduces its own comb-filtering and phase smearing, and that smearing looks superficially like some of the spectral irregularity detection is trying to spot. A heavily widened human mix is a recognisable false-positive shape.

So the useful reading is procedural. If a track reads high and also shows an extreme stereo picture — either near-mono or unusually wide with low correlation — your uncertainty should get wider, not narrower. The picture tells you the measurement conditions were unusual. It does not tell you the answer.

Practical mix checks it answers well

Setting provenance aside, this is a genuinely useful mastering instrument, and the questions it answers cleanly are worth listing.

Is my low end wide enough to cancel on a club system or a phone? Look at side energy in the lowest band; above about 15% there deserves attention. Has a channel been inverted between the mix and the master? Look for a negative overall correlation, or a single band that reads negative while the rest do not. Did the widener go further than I intended? Compare side energy against the 2–15% range and check whether mono-fold loss grew. Is the whole mix leaning to one side? The channel balance figure catches panning slips that ears acclimatise to over a long session. And does the image really open in the chorus, or does it only feel like it? The time trace shows the answer instead of trusting memory.

What this measurement cannot do

Band splitting uses second-order filters with gradual slopes, so per-band figures are indicative rather than surgical — energy near a crossover is shared between neighbours. Correlation is computed across the whole excerpt, so a track that is narrow for two minutes and wide for two minutes reports something in the middle; read the time trace next to it before drawing conclusions. Measurement is bounded to five minutes taken from the middle of the file, which means a long DJ set is described by its middle rather than its whole. And the browser's own decoder reconstructs lossy files, so a heavily compressed MP3 can read slightly differently from the master it came from.

None of that changes with a paid tool. It is the nature of summarising four minutes of moving stereo image in a handful of numbers, which is why the trace and the band table exist alongside the headline figures.

Where it fits in a provenance workflow

Run the detector first and read the range rather than the midpoint. If the estimate is high enough to act on, work through the context checks before you act: file metadata for encoder and software fields, a spectrogram for repeated blocks and edit boundaries, a loudness reading for how hard the master was limited, a stereo reading for how the image was shaped, and a stability check across several excerpts to see whether the track is uniform or patched together.

Each of those explains the score. None of them replaces the two things that actually settle the question: the artist's own account, and the working files. If someone can show you stems, a session, a take history or a credit trail, that is worth more than every measurement on this site combined — and if they cannot, that is a fact about the conversation rather than about the waveform.

The short version

Stereo width, correlation and mono compatibility describe how a mix was placed and processed, not where it came from. They are excellent for catching wide bass, inverted polarity, an over-enthusiastic widener or a panning slip, and they are close to worthless as provenance evidence because a few minutes of processing can produce or destroy any reading you like. Use them to explain a detector result — an extreme image is a false-positive risk, not a confirmation — and leave the actual provenance work to metadata, spectrograms, stability across excerpts, and asking for the session files.

Try the free AI music detector

Frequently asked questions

  • No. Width is set by mixing and mastering decisions made after the music exists, and any track can be widened or narrowed in minutes. Some generators do produce a characteristically even spread, but so does plugin-widened human production, so width alone is not evidence of origin.

More reading