Vocabulary gap widening
A Socratic walk-through of vocabulary gap widening — reasoned out one step at a time, not lectured.
The question we started with
THE QUESTION #Why does a small early gap in vocabulary widen every year instead of closing?
Two five-year-olds arrive at school a few hundred words apart. Both are taught by the same teachers, in the same room, from the same books, for the next decade. The intuition — and the whole hope of universal schooling — is that common treatment should pull them together. Instead the distance between them tends to grow. What is doing that? It is tempting to answer "inequality", but notice how unsatisfying that is: we have specified equal treatment from age five onward, so whatever widens the gap has to be something the equal treatment itself permits.
Reasoning it through
REASONING #Before reaching for any mechanism, do the arithmetic, because a surprising amount of the puzzle dissolves in it. Suppose both children add to their vocabulary at the same proportional rate — fifteen percent a year, say, which is a stand-in, not a measurement. Start one at five thousand words and the other at four thousand, a gap of one thousand. After a year: 5,750 and 4,600. The gap is 1,150. After another: 6,613 and 5,290, a gap of 1,322. Nothing has been added to the unfairness. Each year the gap is simply multiplied by the same 1.15 that everything else was multiplied by, so gap after t years is the original gap times 1.15 to the power t.
That inverts the question. Under multiplicative growth, equal rates do not close an absolute gap — they widen it, automatically, forever. Closing it would require the behind child to grow faster. So the thing needing explanation is not the widening; it is why anyone expected convergence.
But look what the same arithmetic gives with the other hand. Divide the two numbers rather than subtracting them and the ratio is 1.25 in every single year. On a proportional scale, nothing whatever has happened. So part of "the gap widens" is a statement about which scale you chose, and an honest account has to say so before it claims a mechanism.
Now the part that is not arithmetic. Is growth actually multiplicative — does the rate depend on the stock? After the first years, most words are not taught; they are inferred from encountering them in comprehensible surroundings. And inferring a word from context requires understanding the surroundings, which requires knowing the other words. That makes today's stock an input to tomorrow's growth, which is exactly the multiplicative premise.
It is worse than a smooth dependence, though — it is closer to a threshold. Work on lexical coverage suggests a reader needs to know something like 98% of the running words in a text for comfortable unassisted comprehension, with understanding degrading markedly by around 95%. Above the threshold, the unknown words are surrounded by known ones and the context does the teaching for free. Below it, the unknowns cluster, each obscuring the evidence needed to decode the next, and the page yields almost nothing. Identical exposure, radically different yield.
A second amplifier follows: reading below the threshold is unpleasant, so it is done less, and the child self-selects thinner texts. Exposure, which we assumed equal, quietly stops being equal. Stanovich named this family of effects in 1986 after the verse about those who have being given more.
The analogy
THE ANALOGY #Two people learn a card game by watching hands played. One already knows most of the rules, so a single unfamiliar move is surrounded by moves she understands, and she can work out what the new rule must be. The other knows few rules and sees only arbitrary gestures — the same hand teaches him nothing, because the evidence he would need to interpret it is the thing he lacks.
a card game has a finite rulebook and therefore a ceiling where the knowledgeable player stops gaining, whereas vocabulary is effectively open-ended; and the beginner can simply be told the rules, which is a reminder that direct teaching is the intervention this whole mechanism argues for.
Clarifying the model
THE MODEL #Two refinements, and then the honest damage.
First, the arithmetic here is the same arithmetic as compound interest, but the situation is stronger in one respect and weaker in another. Stronger: a savings rate is fixed by contract and does not care how large the balance is, whereas here the rate itself rises with the stock, so the divergence is driven twice over. Weaker: money compounds mechanically, whereas a vocabulary can be intervened upon, and explicit teaching of words is precisely a way to raise the behind child's rate rather than merely her stock.
Second, the famous "thirty million word gap" is not load-bearing here and I would not rely on it. It came from a small study of a few dozen families with hourly observations extrapolated across years, and a later attempt with a broader sample did not reproduce the magnitude. The mechanism above does not need it.
Now let me try to break my own account. It predicts widening, so a measure that converges should falsify it. Decoding converges — children who start behind at sounding out words routinely catch up. But that is what the mechanism predicts too, because decoding has a ceiling: once you can read any word accurately there is nowhere further to go, and ceilings compress differences. Divergence should appear only in open-ended stocks, which is where longitudinal reviews most often find it — for vocabulary and comprehension rather than for decoding. Even there the evidence is genuinely mixed, with plenty of studies finding stability or convergence.
That is the weakest link, and it is a measurement problem more than a data problem: vocabulary test scores are not an interval scale, so a widening in points does not straightforwardly mean a widening in words known. Softest to firmest, then: the fifteen percent rate is illustrative and carries no weight; the coverage thresholds near 95% and 98% are reasonable estimates that vary with text and task; the claim that words are mostly learned from context rather than taught is well supported; and the derivation that equal proportional growth widens an absolute gap is certain, because it is arithmetic.
A picture of it
THE PICTURE #How to readThis is a model, not data — both lines grow at exactly the same fifteen percent a year, so the figure shows what happens with no extra unfairness after year zero. Read the vertical distance between the curves at each year: 1,000 words at the start, about 2,000 by year five. Then read the two lines as a ratio instead and the picture reverses — the upper line is 1.25 times the lower at every point on the axis. The figure is really two claims at once: absolute gaps widen under multiplicative growth, and proportional ones do not, so the scale you pick decides what you see.
What became clearer
WHAT CLEARED #Widening is the default, not the anomaly. Under multiplicative growth an equal rate preserves the ratio and stretches the difference, so no fresh injustice is needed each year to produce a growing gap — and the honest half of that is that the "widening" is partly a choice of scale. The load-bearing claim is the multiplicative premise itself: that the rate of vocabulary growth depends on the vocabulary already held, because words are mostly inferred from context and inference needs the surrounding words. That premise rests on the coverage threshold, which is well motivated and moderately measured, and it is what makes explicit vocabulary teaching — raising the rate, not just the stock — the intervention the model actually recommends.
Where to go next
ONWARD #- Why decoding converges while vocabulary diverges, and what other school measures have ceilings.
- Whether score-scale artefacts explain a meaningful share of the reported widening in longitudinal studies.
Key terms
TERMS #| Term | What it means |
|---|---|
| Matthew effect | Stanovich's name for cumulative advantage in reading, where existing skill accelerates the acquisition of more. |
| Lexical coverage | the proportion of running words in a text that a reader already knows, the variable behind the comprehension threshold. |
| Incidental word learning | acquiring a word's meaning from encountering it in context rather than from being taught it. |
Every term the collection defines is gathered in the glossary.