Marker disagreement
A Socratic walk-through of marker disagreement — reasoned out one step at a time, not lectured.
The question we started with
THE QUESTION #Why do two trained examiners hand the same essay different grades even when both follow the same rubric?
Two examiners are trained together, work from the same rubric, and mark the same script. They hand back different grades. The instinct is that somebody made a mistake, and it suggests a remedy: a better rubric, more training, tighter criteria.
But consider what would have to be true for a rubric to eliminate the difference. It would have to convert every judgement into a determination — so that "develops an argument with insight" is checked the way a spelling is. If a criterion could be applied that mechanically, we would not need an examiner for it. So the disagreement may not be a failure of the rubric at all; it may be a report on what kind of thing is being judged. How would we tell those apart?
Reasoning it through
REASONING #Break the mark down into what could produce it. There is the quality of the script — the thing we want. There is the marker's general severity: some examiners run a mark or two low on everything. There is the interaction between this marker and this script: an examiner who rewards structural discipline meets an essay that is disorganised but original. And there is the occasion: the twentieth script of the evening, the one that followed a brilliant answer.
Now sort those by whether they can be removed, because that sorting is the practical payoff.
Severity is a constant offset, and a constant is correctable: give a marker scripts whose value you already know, find her offset, subtract it. This is what statistical moderation does, and it works — which is why severity, the difference examiners most often worry about, is the least serious problem in the list.
The interaction is a different animal. It is not an offset, because it is not the same for every script: it is this marker's reading of this essay, and a correction that helped the disorganised-but-original script would harm the tidy pedestrian one. And here is the awkward part: much of that interaction is not error. It is two competent readers weighing originality against control differently, on a script where the two genuinely conflict. A rubric can narrow the disagreement by telling them which to weigh more — but only by deciding in advance something the writing itself made a live question.
Occasion effects are the third kind, and the most clearly unwanted. The best-documented is the contrast effect: a script tends to be marked lower when it follows a strong one and higher when it follows a weak one, because the preceding script shifts the marker's working standard. That is straightforwardly noise, invisible to the marker, and no rubric addresses it, because a rubric says nothing about what came before.
So can anything be done about the part that will not subtract? Yes, and it is arithmetic rather than training. If independent marks each correlate at r with the true value, averaging k of them yields a reliability of k times r, divided by one plus k minus one times r — the Spearman-Brown relation. Put a single-marker reliability of 0.6 into it, purely as an illustration and not as a claim about any real examination: two markers give 1.2 over 1.6, which is 0.75; three give 1.8 over 2.2, which is about 0.82; four give 2.4 over 2.8, about 0.86; six give 3.6 over 4.0, exactly 0.90.
Look at the shape of those numbers. The second marker buys 0.15. The third buys 0.07. The sixth buys about 0.02. Independent marks are the only lever that touches the un-subtractable part, and the lever gets heavier and heavier for less and less.
The analogy
THE ANALOGY #Think of a panel of wine judges scoring the same bottle. One judge is a known hard scorer — you can find that out from bottles of known standing and adjust for it. But another simply prefers acidity to weight, and on a wine where the two trade off, his score differs from his colleague's for a reason no calibration removes, because it is not an error about the wine; it is a difference about what a good wine is.
a wine panel is not obliged to agree, and its published score is openly a consensus of tastes, whereas an examination issues a single number as a statement about the candidate — so the same irreducible spread that is honest in one setting becomes an injustice in the other.
Clarifying the model
THE MODEL #test-validity.md treats reliability as a precondition and moves past it to ask whether the inference from a score is sound; the fixed point of difference here is that this piece opens the reliability box on a mark produced by judgement, and finds inside it a component that is not noise at all.
That yields the trade-off examiners actually face. Break an essay into analytic criteria — organisation, evidence, mechanics, each scored separately — and agreement rises, reliably. But it rises partly because you have replaced the judgement of the whole with a sum of narrower judgements, and whatever quality lived in the whole is now unscored. In test-validity.md's terms that is construct under-representation bought with reliability. Holistic marking keeps the construct and pays in disagreement. Neither choice is free, and a rubric is a position on this trade-off rather than an escape from it.
One further caution about the evidence. Reported marker-agreement figures vary enormously between subjects and mark schemes, and with how agreement is computed — exact agreement, adjacent agreement and correlation are three different numbers from the same data. I used a stipulated value above rather than a published one, because a real figure quoted without its subject and its definition would be worse than no figure.
A picture of it
THE PICTURE #How to readEach point is the Spearman-Brown reliability of the average of that many independent marks, computed exactly from a stipulated single-marker reliability of 0.6 — an illustration, not a measurement from any examination. Read the vertical jumps between neighbouring points rather than the heights: the step from one marker to two is by far the largest, the step from two to three is under half of it, and by the seventh the curve is nearly flat. That flattening is the practical result — double marking captures most of what averaging can buy, and no realistic number of markers reaches certainty.
What became clearer
WHAT CLEARED #Two examiners disagreeing are usually doing three different things at once, and only one of them is a mistake. Severity is a constant and can be subtracted. Occasion effects such as contrast are noise and can be diluted by averaging. The marker-by-script interaction is neither: it is what happens when a genuine judgement is required and two qualified readers weigh the criteria differently, and it is the residue a rubric can only redistribute.
The load-bearing claim is that severity is a stable additive trait, because that is what licenses moderation. It has an obvious test: seed the same set of scripts into different markers' batches at different positions, and see whether each marker's offset from the consensus stays the same across batches. If a marker's severity drifts with fatigue, with the batch she happens to be given, or with position in the pile, then it is not a constant, subtracting a marker mean is the wrong correction, and moderation is silently making some scripts worse.
Where to go next
ONWARD #- How many-facet Rasch measurement estimates marker severity and script difficulty at once, and what it assumes to do so.
- Whether comparative judgement — asking markers only which of two scripts is better — escapes the severity problem.
Key terms
TERMS #| Term | What it means |
|---|---|
| Rater severity | a marker's general tendency to score above or below the consensus, treated as an additive offset that moderation can remove. |
| Contrast effect | the shift in a script's mark caused by the standard set by the immediately preceding script. |
| Spearman-Brown formula | the relation giving the reliability of an averaged measurement from the reliability of a single one and the number averaged. |
| Analytic and holistic marking | scoring separate named criteria and summing them, versus judging the whole response at once. |
Every term the collection defines is gathered in the glossary.