Pilot-to-scale failure
A Socratic walk-through of pilot-to-scale failure — reasoned out one step at a time, not lectured.
The question we started with
THE QUESTION #Why does a teaching programme that worked wonderfully in trial fail once every school adopts it?
A programme is trialled in eight schools. The results are strong, the evaluation is competent, the teachers are enthusiastic. It is rolled out to eight hundred schools, and the effect shrinks to something indistinguishable from nothing. The usual verdicts are betrayal or incompetence: the ministry ruined it, the teachers did not follow the manual, the original team oversold.
Sometimes that is true. But the pattern is too regular to be a story about villains — John List named the shrinkage the voltage drop precisely because it recurs across education, health, policing and business. Which suggests the assumption to question is not who failed, but what a pilot was ever measuring.
Reasoning it through
REASONING #Ask a deceptively simple question: the pilot produced a number, but a number describing what?
It describes the effect of that programme, delivered by those people, to those participants, under those conditions, at that size. Every clause is a condition, and the point of scaling is to change most of them — so a pilot estimates a conditional effect, and rollout deliberately alters what it was conditional on. The clauses fail in different ways, so go through them.
Start with the crudest: the number might not be real. A small trial has a wide confidence interval, and programmes are selected for scaling precisely because they looked good. Choose the highest estimate from a set of noisy estimates and you have selected partly on the noise — so the act of promotion itself guarantees the follow-up will disappoint, even when the underlying effect is genuinely positive. Nothing has to go wrong for this one to bite. It is arithmetic.
Then the people delivering it. Pilots run on the designer's attention, hand-picked staff, and the visible fact of being watched. Rollout replaces all three with whoever is already employed, trained in an afternoon. If the effect depended on a scarce ingredient — a charismatic founder, teachers who volunteered — it cannot survive dilution, because there was never enough of that ingredient to go around. The uncomfortable question to ask of any pilot is: was I measuring the programme, or the people who believed in it?
Then the participants. A pilot recruits sites that agree to be studied, and schools that agree are not average schools — their students, resources and appetite for change all differ, usually in the direction that flatters the result.
Then a subtler failure that has nothing to do with fidelity. Some effects work only because they are rare. Train a hundred unemployed people and they compete better for jobs; train a hundred thousand and much of the gain is taken from the untrained, because the number of jobs did not move — a French randomised evaluation of job-placement assistance found exactly this displacement. A credential that distinguishes you when scarce distinguishes nobody when universal. At pilot size the programme is a small perturbation to a system that absorbs it; at full size it is the system.
And finally cost. Some programmes are cheap per pupil only while small, because the pilot borrowed capacity that already existed and was idle. Others get cheaper. Which way it runs is rarely asked before the rollout is budgeted.
So what should we conclude when the effect shrinks? Not, usually, that someone lied. The pilot answered a question honestly. It was simply a narrower question than everyone chose to hear.
The analogy
THE ANALOGY #A recipe developed in a restaurant kitchen, cooked by the chef who invented it, on one plate at a time, is not a claim about what happens when the same dish is produced ten thousand times a night by rotating staff on a fixed budget. The recipe was never wrong. It was tested in a setting whose ingredients — attention per plate, the cook's judgement, produce chosen that morning — are the ones that do not survive volume.
a kitchen's constraints are physical and mostly foreseeable, whereas the sharpest scaling failures are the ones the pilot could not have seen at all — effects that reverse only because the programme becomes common, and which no amount of care at small size would have revealed.
Clarifying the model
THE MODEL #The most useful correction is to the word "failed". If the pilot's effect depended on conditions that scaling removes, then the shrunken result is not a contradiction of the pilot — it is a second measurement, of the thing you can actually deploy, and it is the one that matters. Treating it as new information rather than as a broken promise is what lets an organisation learn instead of assigning blame.
That reframing has a design consequence. A pilot built to produce the most impressive number and a pilot built to predict what rollout will do are different experiments. The second is deliberately unglamorous: ordinary staff rather than the design team, sites that did not volunteer, no founder in the room, and costs that assume no borrowed goodwill. It will report a smaller effect, and that smaller effect is more likely to be the true one.
Two honest limits. This is not an argument that scale always destroys — some interventions get stronger with volume, through network effects or falling unit costs, and the same reasoning predicts that too. And the relative weight of these causes is genuinely contested: distinguishing "the original estimate was inflated by chance" from "the effect was real but the conditions were unreproducible" is difficult after the fact, and both stories fit most cases. What is not contested is that fidelity checklists alone do not solve it, because several of these mechanisms operate at perfect fidelity.
A picture of it
THE PICTURE #How to readThe five boxes are the conditions a result must meet before it survives scaling; the two elements below are the pilot and the rollout, and an arrow means that condition is met there. Read the arrows the pilot is missing first — it cannot vouch for C1, because it is the very estimate in question, nor for C5, because its costs were borrowed. The rollout has the opposite profile: a trustworthy estimate, but the volunteers, the founder's staff and the scarcity that C2 to C4 depended on are all gone. Neither column is a failure; they satisfy different conditions, which is why their numbers differ.
What became clearer
WHAT CLEARED #A pilot does not measure a programme; it measures a programme under conditions that scaling is designed to change. Selection on a noisy estimate, founder attention, volunteer sites and the advantage of being rare are four separate reasons the effect can shrink, and three survive perfect implementation. The practical move is to design pilots that spend their advantages in advance — ordinary staff, assigned sites, honest costs — accepting a smaller number in exchange for one that means something.
Where to go next
ONWARD #- Why the winner's curse makes the most promising trial result the most likely to disappoint.
- How stepped-wedge and staged rollouts turn scaling itself into the experiment.
Key terms
TERMS #| Term | What it means |
|---|---|
| Voltage drop | John List's term for the shrinkage of an effect between a pilot and its scaled version. |
| Displacement effect | a benefit to participants that comes partly at the expense of non-participants competing for the same fixed resource. |
| Selection on the estimate | promoting the intervention with the highest measured effect, which selects partly on chance and guarantees regression on retest. |
Every term the collection defines is gathered in the glossary.