Regression to the mean
A Socratic walk-through of regression to the mean — reasoned out one step at a time, not lectured.
The question we started with
THE QUESTION #Why is an exceptional performance usually followed by a more ordinary one?
An athlete has the season of their life, then a mediocre one. A student sits a test far above their usual mark, then drops back. The clinic's sickest patients, given a new drug, improve. Each of these invites a story — complacency, pressure, the drug working — and the stories are usually the first thing we reach for.
But notice they all have the same shape, across domains with nothing in common. When one pattern turns up everywhere, it is worth asking whether it needs a cause at all, or whether it is a property of measuring twice.
Reasoning it through
REASONING #Galton met this first in the 1880s, looking at the heights of parents and their grown children. Unusually tall parents had children who were tall, but on average less tall — closer to the population's middle. Unusually short parents had children closer to the middle from the other side. He initially read it as a force, something in heredity pulling stock back toward mediocrity, and he named the phenomenon after that reading.
But watch what happens if you run the same data backwards. Select unusually tall children and look at their parents: the parents are, on average, less extreme than the children. A force that pulled generations toward the middle could not also pull backwards in time. So whatever this is, it is not a force.
Then what is it? Try the smallest possible model. Suppose any measured performance has two parts: something durable about the performer, and something that is not — the day, the wind, the marker's mood, which questions came up. Now ask what kind of person ends up at the very top of a single measurement. Almost certainly someone with a high durable part and a favourable draw on the rest, because both are needed to reach the extreme and the extreme is rare enough that you cannot afford to be short on either.
Measure again. The durable part comes back. The other part is redrawn, and there is no reason it should be favourable twice. So the second measurement is, on average, less extreme. Nothing changed about the performer. What changed is that you selected on the total and only half of it repeats.
The general statement is tidier than the story. If two measurements are correlated at r, then the expected second value, measured in standard deviations from the mean, is r times the first. Perfect correlation, no regression; zero correlation, complete regression to the mean; anything in between, partial. That is all it takes — no assumption about causes, mechanisms, or what the measurements are of. Regression to the mean is simply what imperfect correlation looks like when you select on one end.
Which yields the thing worth carrying away: any selection on an extreme value manufactures apparent change. Pick the worst-performing schools for an intervention and they will improve. Pick patients with the highest blood pressure and their pressure will fall. Pick the players who had a career year and they will decline. In every case the effect appears before you have done anything at all.
The analogy
THE ANALOGY #Think of an archer shooting in gusty wind. Where the arrow lands is the aim plus the gust. A bullseye is usually a good aim helped by a lucky moment of calm — so the next arrow, shot with the same aim into a fresh gust, tends to land further out. Nothing in the archer got worse.
The picture suggests two clean components you could separate arrow by arrow, and in real data you never observe the split. Worse, the "gust" is not always random: it can include a bias in the instrument, an illness on the day, a rater's habit. Those are systematic influences, not luck — they simply happen to be uncorrelated with the next occasion, and that is the only property the mathematics needs.
Clarifying the model
THE MODEL #Two misreadings are worth heading off.
The first is that something gets corrected. No individual is pulled anywhere; each performer's next result is drawn from their own distribution as usual. What moves is the average of the group you selected, and it moves because you selected them on a criterion that partly reflected noise.
The second is that variability shrinks over time — that populations march toward uniformity. They do not. Regression is conditional on having selected an extreme, and the same argument run in the other direction produces new extremes from ordinary performers. Galton's own worry evaporates once you notice that the spread of heights was not narrowing generation by generation.
The practical consequence is the important one, and there is a well-known story about it. Kahneman describes flight instructors who observed that trainees praised after an exceptional manoeuvre tended to do worse next time, while those criticised after a poor one improved — and concluded that criticism works and praise backfires. Both observations are exactly what regression predicts if praise and criticism did nothing whatever. Worse, the instructors were being reinforced for the wrong lesson every day, because the pattern is perfectly reliable.
That is why controls exist. If you measure a selected group before and after an intervention, regression and the intervention are hopelessly confounded — the before-after difference contains both, and you cannot tell them apart. A control group selected the same way regresses by the same amount, so the difference between the groups isolates what the treatment actually did. The randomised trial is not just guarding against enthusiasm; it is guarding against arithmetic that produces a result for free.
A picture of it
THE PICTURE #How to readThe steeper line is the naive expectation — whatever you scored, expect the same again. The flatter line is the actual expected second score, and its flatness is the whole phenomenon. Pick any first score on the horizontal axis and read up to both lines: the gap between them is how badly you would be surprised. At the mean of 100 the lines meet and there is nothing to explain; the further out you select, the larger the apparent decline or improvement you will observe with no cause behind it. Halve the correlation and the flat line flattens further; push it to 1 and the two lines merge.
What became clearer
WHAT CLEARED #Regression to the mean is not a phenomenon of athletes, students, or patients — it is a phenomenon of measuring the same thing twice with anything less than perfect agreement, and then choosing what to look at from the first measurement. It requires no cause, which is precisely why it so readily attracts one. And that is its real danger: it hands you a reliable, repeatable pattern with a plausible explanation attached, in exactly the situations — struggling schools, sick patients, bad quarters — where someone is about to claim credit for it.
Where to go next
ONWARD #- How regression to the mean interacts with measurement reliability, and why unreliable tests regress hardest.
- Why difference-in-differences and pre-registered controls are built the way they are.
Key terms
TERMS #| Term | What it means |
|---|---|
| Regression to the mean | the tendency of a second measurement to lie nearer the mean than a first extreme one, whenever the two are imperfectly correlated. |
| Correlation coefficient (r) | the factor by which a standardised score is expected to shrink from one measurement to the next. |
| Selection on an extreme | choosing cases by their value on one measurement, which is what triggers the effect. |
| Control group | a comparison group selected the same way, so that regression affects it equally and cancels out of the comparison. |
Every term the collection defines is gathered in the glossary.