Pre-registered analysis
A Socratic walk-through of pre-registered analysis — reasoned out one step at a time, not lectured.
The question we started with
THE QUESTION #Why does deciding your test in advance change what the very same number means afterwards?
Two laboratories run the same experiment, collect the same measurements, and both report p = 0.03. One of them lodged its analysis plan, publicly and with a timestamp, before a single participant was recruited. The other decided how to analyse things once the data were in.
Most people's instinct is that the second lab was less disciplined but that the number itself is the number — arithmetic does not care who typed it. So here is the assumption worth pressing on: is a p-value a property of the data alone?
Reasoning it through
REASONING #Ask what a p-value actually claims. It is a statement about a procedure: if there were no real effect, and I ran this test, the chance of getting a result at least this extreme is 3 in 100. Every word of that sentence except "no real effect" describes something the analyst does, not something the data are.
So let us test whether the procedure was really the same in both labs. In the second lab, the analyst had choices. Should the two outcome measures be reported separately or combined? Should the eleven participants who misunderstood an instruction be excluded? Should age be a covariate? Should collection have stopped at 60 or continued to 80?
Ask yourself: how many of those decisions would have been made differently had the first look at the data been discouraging? Not dishonestly — just plausibly. The exclusion looks principled when it tidies a result and unnecessary when it does not.
Now count. If each of four such decisions has two defensible settings, there are sixteen analyses that could have been reported. Suppose each has roughly the nominal 5 percent chance of a false positive, and the analyst reports whichever succeeds. The chance that at least one of sixteen lands under 0.05 is far above 5 percent. The reported number still says "3 in 100", but the procedure that generated it was "search sixteen and report the winner", and that procedure's false-positive rate is not 3 in 100.
This is not speculation. Simmons, Nelson and Simonsohn simulated exactly this in 2011 with four ordinary flexibilities — two dependent variables, adding observations, controlling for gender, dropping a condition — and found the false-positive rate climbing from the nominal 5 percent to about 61 percent when all four were available.
Now the subtle part, and it is the part that makes pre-registration necessary rather than merely tidy. Gelman and Loken argued in 2013 that the analyst need not consciously try several analyses at all. They only need to have chosen this one because of what the data looked like. One path is walked; the other fifteen were still live at the moment of choosing, and they are what the p-value's guarantee was computed against. They called it the garden of forking paths, and it means that a sincere researcher who ran exactly one test can still produce a number whose stated error rate is fiction.
Which brings us back to the commitment. What does lodging the plan actually accomplish? It does not improve the data, the instrument, or the arithmetic. It removes the branches. Once the plan is timestamped, the analysis is no longer a choice made in the presence of the data — it is an obligation contracted in their absence. The set of analyses the number could have come from collapses to one, and only then does the sentence "3 in 100 if there is no effect" describe what actually happened.
Notice the direction of the work: the commitment costs the researcher freedom and buys the reader the right to take the number at face value. That is the whole trade.
The analogy
THE ANALOGY #It is calling your pocket in pool. You point at the far corner, then strike; the ball drops. Say nothing beforehand and sink the identical ball into the identical pocket, and you have shown you can hit balls into pockets — which of the six it was gets decided after the fact, by whichever one it went into. Same cue, same ball, same trajectory, entirely different claim supported, and the difference lives wholly in what you gave up the option of saying.
at a pool table everyone can see there are six pockets, whereas in an analysis the number of alternative pockets is invisible, unbounded, and — per Gelman and Loken — often unknown even to the person taking the shot.
Clarifying the model
THE MODEL #Pre-registration is often misheard as a ban on exploration. It is not. It is a label. An analysis specified in advance carries the error guarantee it advertises; anything else is exploratory and reported as such, generating hypotheses rather than testing them. Both are legitimate science; conflating them is what breaks.
Nor does it make a result true. A pre-registered study can be underpowered, badly measured, or simply unlucky. What commitment protects is one specific thing: the correspondence between the stated error rate and the real one.
The strongest form goes further. A Registered Report submits the plan to peer review before data collection and, if accepted in principle, the journal commits to publishing regardless of outcome — which removes the other selection filter, the one operating on which finished studies ever appear.
Is there evidence it bites? Kaplan and Irvin reported in 2015 that among large NHLBI-funded cardiovascular trials, 57 percent published before 2000 showed significant benefit for the treatment, against 8 percent after — the change coinciding with mandatory registration on ClinicalTrials.gov. That is observational, other practices changed over the same period, and it should be read as strongly suggestive rather than decisive.
A picture of it
THE PICTURE #How to readthe flat line is the error rate every one of these studies would report, and the bars are the rate actually incurred as ordinary undeclared choices are added, using the simulated values from Simmons, Nelson and Simonsohn (2011); pre-registration is what pins the bar back onto the line.
What became clearer
WHAT CLEARED #A p-value is not a property of a dataset but of the procedure that produced it, and a procedure includes every choice that was still open when the data were in view. Committing in advance does not change the number — it changes what the number is entitled to claim, by destroying the alternatives it would otherwise have been quietly selected from.
Where to go next
ONWARD #- Multiverse and specification-curve analyses, which report all the defensible paths instead of committing to one.
- How pre-registration interacts with adaptive and sequential trial designs, where stopping rules are planned but flexible.
- Whether deviations from a registered plan can be handled honestly, and what a good deviation statement looks like.
Key terms
TERMS #| Term | What it means |
|---|---|
| p-value | the probability of a result at least this extreme, assuming no real effect and this procedure. |
| Researcher degrees of freedom | the defensible analytic choices left open once data are in hand. |
| Garden of forking paths | Gelman and Loken's term for the paths not walked that still inflate the true error rate. |
| Registered Report | a format where the plan is peer-reviewed and accepted before data collection. |
Every term the collection defines is gathered in the glossary.