THIS EXPLANATION
THE ROOM
EDU·32 Education & Learning 6 MIN · 8 STATIONS

Teaching to the test

A Socratic walk-through of teaching to the test — reasoned out one step at a time, not lectured.

abcdefgh
a

The question we started with

THE QUESTION #

Why does a school measure stop meaning what it once did as soon as everyone is judged by it?

A reading test is chosen for an accountability system precisely because it works. Before anyone was judged by it, its scores tracked what inspectors, teachers and later employers independently thought about those children. Then it is made consequential — and within a few years the scores rise while nobody outside the system can see the improvement. The usual explanation is cheating, and cheating exists, but it is far too rare to account for this. So the assumption worth questioning is subtler: that the measure's usefulness was ever a property of the measure at all.

b

Reasoning it through

REASONING #

Start with what a test is. It is a sample from a domain — a few dozen items standing in for everything a child ought to be able to do with written language. Why should a sample tell you about the whole? Because of an assumption nobody states: that the school's effort was not allocated according to which items would appear. Effort was spread across the domain for its own reasons, so a sample from the domain represented it.

Now attach stakes and ask a school's honest question: what is the cheapest way to raise this number? Teaching the whole domain better is expensive and slow. Teaching only the sampled part and dropping the rest is cheap. Drilling this paper's formats and conventions is cheaper. Concentrating on children just below the reporting threshold — where a small gain moves the published figure and the same effort on a child far below it moves nothing — is cheapest of all. Below that lie the improper options: exclusions, reclassifications, cheating.

Notice what happens from the second option down. Each raises the score without raising the thing the score stood for, and none requires bad faith; a head teacher choosing to spend the year on what the children will be examined on is doing an obviously defensible thing.

So the correlation between score and domain was never structural. It was an artifact of the effort allocation that held while nobody was optimising the score, and the stakes did not corrupt the measure — they removed the condition that made the sample representative.

Can we test that? First direction: if stakes are doing the work, the same children measured on an unrelated low-stakes instrument should not show the same gains. That is what is found. The pattern was visible from the beginning — John Jacob Cannell's 1987 survey found essentially every US state reporting its pupils above the national norm, an arithmetical impossibility that became known as the Lake Wobegon effect — and later comparisons of high-stakes state tests against the low-stakes national assessment sat by the same cohorts show state gains repeatedly outrunning the independent measure.

Now the other direction, where the account gets sharper. Does attaching stakes corrupt every measure? No. Imagine a test that sampled the whole domain unpredictably — any item, no forewarning of which. Then the only way to raise the score is to teach all of it, and "teaching to the test" is simply teaching. Corruption is caused not by stakes alone but by stakes acting on the gap between proxy and target: stakes are the pressure, the slack is the opening. So the invariant is: decay is proportional to the slack between what is measured and what is meant, and stakes only convert that slack into behaviour. Which is why the remedy is never "remove stakes" but shrink the slack or randomise the sample — both expensive, which is why cheap measures keep being chosen.

c

The analogy

THE ANALOGY #
THE FIGURE

Think of a landlord judged by an inspector who, it turns out, only ever opens the front room. The front room becomes immaculate. The inspection was genuinely informative once, because nobody knew which room would be opened and the front room was as good a sample of the house as any other.

WHERE IT BREAKS DOWN

rooms are separable, so the landlord can neglect the back of the house at no cost to the front, whereas school subjects share underlying capabilities — narrowing onto tested reading does raise some untested performance too — which is why the corruption is partial rather than total, and why arguing about it is harder than it looks.

d

Clarifying the model

THE MODEL #

Three refinements, and one of them is a genuine objection.

First, the shape of the metric matters as much as its content. A measure defined as "percentage above a cutoff" makes a child a point below the line far more valuable than one well below or comfortably above, and the resulting triage responds to the formula rather than the test. Change to a value-added or whole-distribution measure and the gaming does not vanish; it relocates.

Second, this is the dynamic twin of a static problem. A validity argument asks whether a score supports a claim, and finds the tasks may sample only a corner of what was meant — a gap present on day one, with no stakes anywhere. Accountability adds a force that widens that same gap year on year, because the corner is now worth shrinking. Same gap, examined at rest and under load. Both differ again from grade inflation, where the taught content is unchanged and the awarder becomes more generous: three mechanisms, one symptom of a signal losing its meaning.

Third, the objection. Whether narrowing is harmful is contested. If the tested content is the content that matters most — basic literacy and number for children previously getting neither — then reallocating time toward it is a gain, not a corruption, and some evaluations of accountability regimes do find real improvements on independent measures, particularly in early mathematics. "Teaching to the test" corrupts exactly to the extent that the untested remainder mattered, which is a judgement about curriculum, not a fact about measurement.

e

A picture of it

THE PICTURE #
Teaching to the test
Teaching to the test Each spoke is a slice of what a school could spend time on, and each closed shape is how one year's effort was distributed across them -- a schematic of the pattern, not measured data from any particular school. Compare the two shapes rather than any single point: the total area is roughly unchanged, because the school gained no hours, but it has been pulled toward the two spokes the measure can see and away from the three it cannot. The score genuinely rises, and it rises by redistribution rather than growth -- exactly the thing a score cannot report about itself. {"generator":"[email protected]","source":"../Socrates/.diagram-cache/_src/teaching-to-the-test.md","sourceIndex":1,"sourceLine":4,"sourceHash":"8367f13f317538bffbf0f6e8bc8cb61b1e5e2d7336a75c7fa01a99dcd6e67695","diagramType":"radar","layoutVariant":"source","repairedDuplicateIds":[],"motion":"entrance-with-reduced-motion-fallback","presentation":"editorial","attempt":1,"viewBox":{"x":0,"y":0,"width":968,"height":767},"qa":{"passed":true,"findings":[]}} Tested item types Test format drill Untested topics Writing at length Science and arts time Before the measure carried stakes After the measure carried stakes

How to readEach spoke is a slice of what a school could spend time on, and each closed shape is how one year's effort was distributed across them — a schematic of the pattern, not measured data from any particular school. Compare the two shapes rather than any single point: the total area is roughly unchanged, because the school gained no hours, but it has been pulled toward the two spokes the measure can see and away from the three it cannot. The score genuinely rises, and it rises by redistribution rather than growth — exactly the thing a score cannot report about itself.

f

What became clearer

WHAT CLEARED #
WHAT CLEARED

A test measures a domain only while nobody is aiming at the test. Its usefulness was never a property of the instrument but of the relationship between the instrument and behaviour not yet optimising against it. Attaching consequences does not damage the items — it makes the cheapest route to a higher score one that passes outside the domain, through the sampled corner, the format, and the children nearest the threshold. And because decay is proportional to the gap between proxy and target, the only real fixes close that gap or make the sample unpredictable; everything else moves the gaming somewhere less visible.

g

Where to go next

ONWARD #
  • How value-added models change what is worth gaming, and what new distortions they introduce.
  • Whether inspection judgements, being harder to predict, decay more slowly than test scores do.
h

Key terms

TERMS #
TermWhat it means
Goodhart's lawa measure ceases to be a good measure once it becomes a target.
Score inflationa rise in scores on one test not matched on other measures of the same domain.
Lake Wobegon effectthe finding that nearly all US states reported above-average pupils, named for the fictional town where every child is above average.
Educational triageconcentrating effort on pupils near a reporting threshold, where it moves the published figure most.

Every term the collection defines is gathered in the glossary.

Nearby on the shelf

4