Advice from top students
A Socratic walk-through of advice from top students — reasoned out one step at a time, not lectured.
The question we started with
THE QUESTION #Why does the study routine of the highest scorer so often fail everyone who copies it?
Every cohort produces one. The student who came top is asked how she did it, and she answers honestly: she never used flashcards, she worked in ninety-minute blocks, she skipped the seminar readings and went to the primary sources instead. The advice is written down, passed around, and adopted — and a term later most of the people who adopted it are exactly where they were.
The natural explanation is that they did not follow it properly. But a stranger possibility deserves trying first: that the advice was never evidence about what works, however sincerely given. Not because she lied, but because of how she came to be the person we asked.
Reasoning it through
REASONING #Ask what we would need in order to learn, from observation, whether a routine causes good results.
We would need to see the routine used by many people and the outcomes that followed — including the outcomes that followed when it did not work. That is the whole content of a causal claim: not "this preceded success" but "success is more likely with it than without it". Which means we need two comparisons, the routine's failures and the alternatives' successes.
Now look at what the classroom actually hands us. We do not sample students and then read off their routines. We sample on the score, take the person at the top, and ask her afterwards what she did. Every question about the routine is asked of someone already selected for the outcome.
Follow the consequence. Suppose thirty students independently invented her routine and it does nothing at all. Some of them still finish near the top, because scores also depend on prior knowledge, on sleep, on which questions came up. One of those becomes the person we interview. The other twenty-nine, who used the identical routine and finished in the middle, are never asked anything — there is no ceremony for them. So a routine with no effect whatsoever still generates a confident testimonial, reliably, every year. That is the core of it: the interview is guaranteed to occur regardless of whether the routine works, so hearing it carries almost no information.
Push a little further, because there is a second and separate problem hiding behind the first. Suppose the routine did contribute. Did it contribute because it is a good routine, or because it suits her? A student with strong prior knowledge can skip the seminar readings and go to the sources, and be right to. A student without it will drown there — for him the reading was the scaffolding, not the shortcut. The routine that is optimal at the top of a distribution is frequently the wrong one lower down, and this is not a failure of effort. It is what it means for a treatment to interact with the person receiving it.
And a third, quieter one. What she reports is what she noticed, and people are poor witnesses to their own habits. The inputs most likely to go unmentioned are the ones that were always there: a parent who is a chemist, twelve years of reading for pleasure, a tutor, a quiet room. She is not concealing them. They are simply not experienced as part of the routine, so the story we get systematically omits the advantages and keeps the visible techniques.
Which puts the odd conclusion in view: how far this advice spreads is nearly unrelated to how well the routine works, and quite strongly related to how unusual it sounds. "I did the practice papers, spaced over eight weeks" attracts no attention. A striking answer travels. The selection happens twice — first on the score, then on quotability — and both filters run on something other than truth.
What would carry information instead? Sampling on the routine rather than the outcome: take everyone who used it and see where they landed, including the ones who landed nowhere. That is precisely what the experimental literature on study technique does, and it is why its findings differ so sharply from the testimonial ones. The methods with the strongest evidence — retrieval practice, spacing, mixing problem types — are unglamorous, well replicated, and almost never what the top student volunteers, partly because they are boring and partly because they feel harder while working better.
One caution against over-correcting: this is an argument about what the observation can support, not proof that she is wrong. Her routine may be excellent. The claim is narrow — her success gives us close to no reason to believe it, so the routine must be judged on its mechanism and on whether it holds up when tested.
The analogy
THE ANALOGY #Picture a room of lottery winners, each explaining their system — birthdays, a dream, numbers that had not come up in a while. Every account is sincere, every one is followed by a genuine win, and every one is worthless, because the room was assembled by the outcome. Somewhere outside it are the millions who used those same systems and won nothing, and the room has no door for them.
a lottery is pure chance while an exam is substantially skill, so the analogy overstates the futility — study routines really do move scores, which is exactly why the testimonial is tempting, and why the fix is better evidence rather than giving up on the question.
Clarifying the model
THE MODEL #Three refinements.
First, survivorship bias is usually explained as data that went missing. Here nothing is missing — the middling students are in the room, we simply never point the question at them. The bias is in the sampling rule, not in a gap in the records, and that is why it survives having complete data.
Second, this does not make top students bad informants about everything. They are often very good on what a question demands, what the marking rewards, or how a topic hangs together — claims we can check directly against the work. It is the causal claim, this is what produced the score, that the selection specifically destroys.
Third, the same reasoning applies to any advice collected from winners — founders, athletes, authors describing their mornings. Whenever the sample was assembled by the outcome, the interesting question is the one nobody asked: how many did the same thing and are not in the room.
A picture of it
THE PICTURE #How to readLeft is what students did, right is where they finished, and the widths are illustrative rather than measured. The advice you hear comes entirely from the thin band running into Top scores. Compare it with the far wider bands leaving the same routine for the middle — identical method, never interviewed. Then notice that other routines reach the top too. Judging the routine means comparing the proportions leaving each left-hand node, which is the one comparison a testimonial cannot show you.
What became clearer
WHAT CLEARED #The top scorer's routine fails its imitators because her success is what caused us to ask, so the testimonial would exist whether or not the routine did anything. On top of that, the method best suited to someone at the top is often wrong for someone lacking the advantages she never thought to mention. Sampling on the outcome cannot answer a question about causes — and the fix is not a more articulate winner but a look at everyone who tried it.
Where to go next
ONWARD #- Why retrieval practice and spacing feel less effective than they are while you are doing them.
- Heterogeneous treatment effects: when an average finding still gives the wrong advice to an individual.
- Whether asking the most improved student rather than the top one dodges the problem or moves it.
Key terms
TERMS #| Term | What it means |
|---|---|
| Selection on the outcome | assembling a sample using the result you are trying to explain, which makes it uninformative about causes. |
| Survivorship bias | reasoning only from cases that reached a visible endpoint, with no view of those that did not. |
| Retrieval practice | recalling material from memory rather than rereading it; among the best-replicated study findings. |
Every term the collection defines is gathered in the glossary.