Train-the-trainer cascades
A Socratic walk-through of train-the-trainer cascades — reasoned out one step at a time, not lectured.
The question we started with
THE QUESTION #Why can training ten teachers reach a whole country faster than training a thousand?
A ministry wants every one of its four hundred thousand teachers to learn a new way of teaching reading. It has one team of expert trainers. Somebody proposes the obvious plan: send the experts out to train teachers, region by region, as fast as they can travel. Somebody else proposes training ten people — and claims that will be faster.
On the face of it that is absurd. Ten is fewer than a thousand, and fewer people trained is less progress. But the claim is made constantly, by health ministries and literacy programmes and safety regulators, and they are not all fools. So the arithmetic must be doing something we have not looked at yet. What is it?
Reasoning it through
REASONING #Let us be careful about what we are counting. In the direct plan, the expert team's output per month is a fixed number of teachers — say they can run workshops for two hundred. Next month, two hundred more. The total climbs by the same amount every month, forever. Draw it and you get a straight line, and its slope is set by one thing: the size of the expert team, which is the resource nobody can conjure more of.
Now change one assumption, and only one. Suppose a teacher, once trained, can train other teachers. Notice what that does. The quantity being produced — trained people — is now also the thing that does the producing. The output feeds back into the capacity. That single feedback is what separates a line from a curve, and it is worth pausing on, because it is the whole mechanism and everything else in this piece is a qualification of it.
Follow it through concretely. One master team trains ten. Those ten each train ten, so after the second round there are a hundred trained trainers, plus the original ten. Each of those hundred trains ten, and now there are a thousand. Round four gives ten thousand; round seven, ten million. The pool multiplies by ten per round rather than growing by two hundred per month, so the total is governed by ten to the power of the round number. What you are buying with the cascade is not more effort per round — it is an exponent.
That is also why the comparison feels wrong at the start. In the first round the cascade has trained ten people and the direct plan has trained two hundred, so the direct plan is comfortably ahead and looks obviously right. The lines cross later. Ask yourself what a manager sees three months in, and you can see immediately why cascades get cancelled just before they would have paid off.
Now the honest half. If growth were the only consideration everyone would cascade everything, and the record is much messier than that. The reason is that the thing being copied is not a number, it is a practice, and each round copies it imperfectly. A trainer who has spent three days with an expert can pass on perhaps most of what matters — but only what they themselves understood, and their trainees inherit that reduced version and reduce it again. Multiply a fidelity of, say, four-fifths across four rounds and about two-fifths of the original survives. The same exponent that multiplies reach also multiplies loss. Researchers of cascade training in language education have documented exactly this dilution, along with a characteristic drift by which the later rounds turn a method into a set of slogans about the method.
So the two effects have the same shape, which is why the design question is never whether to cascade but how to hold fidelity up while the numbers run. The measures that help are unglamorous: scripted materials and video so that each round transmits an artefact rather than a memory; monitoring that samples the last round, not the first; keeping the cascade shallow, two or three layers rather than six; and selecting trainers for their ability to teach adults rather than for seniority. Note what these have in common — each one attacks the per-round fidelity, because that is the term sitting in the exponent.
The analogy
THE ANALOGY #Think of photocopying a page. Copying the original a thousand times is slow but every copy is faithful, and the machine is the bottleneck. Copying a copy, then copying that, reaches a thousand pages in three passes — and each generation is a little greyer, until the tenth generation is a smudge that is technically a page.
a photocopier degrades blindly and identically every time, whereas a trained teacher can repair what they received by consulting the original materials or asking a question, which is precisely why written scripts and monitoring change a cascade's outcome and no amount of care changes a photocopier's.
Clarifying the model
THE MODEL #The refinement that connects the steps is that a cascade is not a way of doing more work. The total number of training days delivered is roughly the same either way; what changes is who delivers them and therefore how many can be delivered in parallel. The cascade converts a scarce, non-reproducible resource — expert time — into an abundant one, by making trainees into producers. That is why it is the standard answer whenever demand is national and expertise is thin, and why it is pointless when expertise is plentiful.
The misconception worth correcting is that fidelity loss is a failure of effort, something that better-motivated trainers would avoid. It is structural. Every transmission of a complex practice through a person loses something, so the loss compounds by the same logic as the growth. Treating it as a discipline problem produces exhortation; treating it as an exponent produces the two fixes that actually work — fewer layers, and stronger artefacts carried between them.
One thing to hold loosely: how much fidelity is lost per round is not a constant of nature, and reported outcomes for cascade programmes vary from strong to negligible. The mathematics of the growth is certain; the value of what arrives at the bottom is an empirical question that has to be measured in each programme rather than assumed.
A picture of it
THE PICTURE #How to readread left to right, one round at a time, and watch the entry gain a digit at every step rather than a fixed increment — that digit is the whole argument. Then read it a second time asking what fraction of the original practice is still intact by the rightmost entry.
What became clearer
WHAT CLEARED #Training ten people is faster than training a thousand only because the ten are not the destination — they are the multiplier. The moment output can be fed back into capacity, growth stops being a line and becomes a power, and the scarce expert time stops being the ceiling. The same structure guarantees that whatever the first round failed to transmit is also raised to that power, which is why a cascade is judged not by how many were reached but by what reached them.
Where to go next
ONWARD #- How epidemiologists' reproduction number describes the identical structure with the opposite sign.
- Whether recorded or online instruction removes the fidelity problem or just relocates it.
- Why some programmes deliberately stop at two layers even when they could go deeper.
Key terms
TERMS #| Term | What it means |
|---|---|
| Cascade training | a design in which each cohort of trainees becomes the trainers of the next, in successive rounds. |
| Fidelity | the share of the original practice that survives one act of transmission, and therefore the term that compounds across rounds. |
| Master trainer | the original expert cohort, whose scarcity is the constraint the cascade exists to escape. |
Every term the collection defines is gathered in the glossary.