THIS EXPLANATION
THE ROOM
EDU·08 Education & Learning 7 MIN · 8 STATIONS

Exam equating

A Socratic walk-through of exam equating — reasoned out one step at a time, not lectured.

abcdefgh
a

The question we started with

THE QUESTION #

Why can this year's candidates sit an easier paper than last year's and still be judged on the same scale?

A new paper is set each year, because reusing last year's would be absurd once it has been seen. But a new paper is never exactly as hard as the old one. So a candidate scoring 52 this year and one scoring 48 last year may have done equally well, or not — and there is no telling by looking at the papers, since "how hard is this question" is what nobody can judge reliably in advance.

The obvious fix is to fix the pass mark to a proportion of candidates. But that punishes a strong year and rewards a weak one, which is the injustice the apparatus exists to prevent. So what could tell us whether this year's higher marks mean better candidates or an easier paper?

b

Reasoning it through

REASONING #

Write down what a raw mark contains. It reflects two things at once: how able the candidates were, and how hard the paper was. One observation, two unknowns. That is not a hard problem — it is an unsolvable one, in the strict sense. No care with the marking helps, because the confusion is not in the marking.

Which tells us what to find: a second observation in which one of the two unknowns is pinned down. Not the candidates — they are different people, that is the point. But the questions can be. Suppose a set of items is sat by both cohorts: identical questions, wording, marking. Their difficulty is the same on both occasions, by construction. So any difference in performance on that set cannot be a difficulty difference. It can only be a candidate difference.

That is the whole idea. The common set is called an anchor, and it converts an unsolvable problem into an ordinary one: the anchor measures the cohorts against each other, and once the cohort difference is known it can be subtracted from the total, leaving difficulty as the remainder.

Work it through. Both cohorts sit a 20-mark anchor and an 80-mark paper that changes yearly. Last year: anchor mean 12.0, paper mean 48.0. This year: 12.8 and 52.0.

This year's cohort scored 0.8 marks more on identical questions, which is 0.8/20 = 4% of the anchor scale. Take that as their ability advantage. Applied to the 80-mark paper, a cohort 4% stronger would be expected to score 0.04 x 80 = 3.2 marks more with no change in difficulty. The observed rise is 4.0. So 3.2 marks of it is a better cohort and 4.0 - 3.2 = 0.8 marks is an easier paper. If last year's pass mark was 40, this year's must be 40.8 — candidates need slightly more, precisely because the paper gave it slightly more cheaply.

Notice the structure of that inference. Nobody judged the questions' difficulty; it was deduced from a discrepancy, using the anchor as a fixed point. And notice where the argument's whole weight sits: on the anchor items meaning the same thing on both occasions.

c

The analogy

THE ANALOGY #
THE FIGURE

It is like weighing two sacks of grain on two sets of scales, neither of which you trust. You cannot compare the readings. But put the same iron block on both first: it has not changed weight between weighings, so any difference in what the scales say about it is the scales disagreeing — and once you know how they disagree you can correct their readings of the grain.

WHERE IT BREAKS DOWN

the block's weight is guaranteed by physics and needs no assumption, whereas an anchor item's difficulty is guaranteed only by nobody having seen it, taught to it, or changed the syllabus around it — so the fixed point in an exam is a claim about the world rather than a fact about the object. And a block can be re-checked against a standard; a leaked item cannot be un-leaked.

d

Clarifying the model

THE MODEL #

Four refinements, from the technical to the uncomfortable.

First, the arithmetic above is a deliberate simplification. Real equating does not scale a mean proportionally: the common methods either match the two score distributions percentile by percentile, or fit a model in which each item and each candidate has an estimated parameter and the anchor links the administrations onto one scale. Both do the same conceptual job as my sum, with far more care about the distribution's shape, and neither assumes ability translates linearly from a 20-mark set to an 80-mark one.

Second, the anchor must be representative as well as common: a set drawn from one topic measures the cohort difference on that topic rather than overall, and the whole scale inherits the distortion.

Third — the real vulnerability — the failure mode is silent. Suppose an anchor item leaks, or the syllabus moves so a topic is taught a year earlier. This year's cohort does better on the anchor for a reason unrelated to their ability. The method reads that as a stronger cohort, concludes the paper must have been harder than it was, and moves the pass mark the wrong way. Nothing in the output looks unusual: no error, no anomaly, just a year of grades quietly wrong. That is why anchor security is the operational obsession it is.

Fourth, equating cannot repair a change of construct. If the syllabus is revised so the two papers examine different things, no statistical linking makes the scores comparable, because they are not measurements of the same quantity; the right response is to say the series is broken. Worth adding, since "the algorithm decided the grades" now conflates several operations: the withdrawn model applied to English school grades in 2020, when no exams were sat, was not equating — it standardised teacher assessments against historical school-level patterns, with no anchor and nothing pinning candidate ability down.

How would we detect the silent failure? Not from the equated scores, which look fine. From the anchor's internal behaviour: stable anchor items should retain their difficulty ordering across administrations, so an item that has moved sharply relative to its neighbours — differential item functioning between the two sittings — is the signature of contamination. That is the refuting observation, and honestly stated it detects some breakdowns, not all: a leak affecting the whole anchor uniformly would shift every item together and leave the ordering intact.

e

A picture of it

THE PICTURE #
Exam equating
Exam equating The two axes are the unknowns hidden inside one number. Compare the two points labelled "same mark": they sit at opposite corners and produce an identical raw score, the ambiguity no careful marking can resolve -- a raw mark fixes only your position along that diagonal, never where on it you sit. The anchor supplies the horizontal coordinate independently, which is how "Last year" and "This year, equated" get placed at all. The fifth point is the failure: a contaminated anchor overstates the cohort, sliding the estimate right, and the method compensates by inferring a harder paper than was set. {"generator":"[email protected]","source":"../Socrates/.diagram-cache/_src/exam-equating.md","sourceIndex":1,"sourceLine":4,"sourceHash":"7c6149f838cbf40ad9e132c59aa518ec52cf8e3e09f247d8b4c0402d83504319","diagramType":"quadrantChart","layoutVariant":"source","repairedDuplicateIds":[],"motion":"entrance-with-reduced-motion-fallback","presentation":"editorial","attempt":1,"viewBox":{"x":0,"y":0,"width":720,"height":621},"qa":{"passed":true,"findings":[]}} Marks rise twice over Q1 Paper flatters them Q2 Marks fall twice over Q3 Ability shows through Q4 If the anchor leaks Same mark, hard paper Same mark, easy paper This year, equated Last year Weaker cohort Stronger cohort Harder paper Easier paper What a raw mark cannot separate

How to readThe two axes are the unknowns hidden inside one number. Compare the two points labelled "same mark": they sit at opposite corners and produce an identical raw score, the ambiguity no careful marking can resolve — a raw mark fixes only your position along that diagonal, never where on it you sit. The anchor supplies the horizontal coordinate independently, which is how "Last year" and "This year, equated" get placed at all. The fifth point is the failure: a contaminated anchor overstates the cohort, sliding the estimate right, and the method compensates by inferring a harder paper than was set.

f

What became clearer

WHAT CLEARED #
WHAT CLEARED

Comparability across years is not achieved by writing papers of equal difficulty, which nobody can do, nor by fixing the proportion who pass, which is unjust. It comes from building a constant into the design — questions both cohorts meet unchanged — so one of the two unknowns inside a raw mark can be measured separately and removed. The method's strength and its fragility are one fact: everything rests on that constant genuinely being constant, and when it is not, the error is invisible in the output and must be hunted for in the anchor itself.

g

Where to go next

ONWARD #
  • What a system does when the syllabus changes enough that no honest anchor exists.
h

Key terms

TERMS #
TermWhat it means
Equatingplacing scores from different forms of a test onto a common scale, so a given score means the same thing.
Anchor itemsquestions common to two administrations, whose unchanged difficulty lets the cohorts be compared.
Item response theorymodels estimating a parameter for each item and each candidate, letting forms be linked through shared items.
Differential item functioningan item behaving differently for two groups matched on ability; here, the signal that an anchor has stopped being stable.

Every term the collection defines is gathered in the glossary.

Nearby on the shelf

4