THIS EXPLANATION
THE ROOM
EDU·01 Education & Learning 6 MIN · 8 STATIONS

Ability grouping

A Socratic walk-through of ability grouping — reasoned out one step at a time, not lectured.

abcdefgh
a

The question we started with

THE QUESTION #

Why can sorting students by ability raise one group's results while lowering another's?

The case for sorting students by attainment is almost too obvious to argue: a lesson pitched at the middle of a wide class is too fast for some and too slow for others, and narrowing the spread should let a teacher aim properly. If that were the whole mechanism, sorting would help everyone somewhat and harm nobody. Yet the research does not show that. It shows something lopsided — and it is worth working out why an intervention that sounds symmetric produces asymmetric results.

b

Reasoning it through

REASONING #

Begin with the one thing sorting definitely does: it reduces the spread of prior attainment inside a room. What follows from that alone? A teacher can pitch closer to where the students actually are. Call this the instructional match, and note that it is the entire theory of benefit. If nothing else changed, both groups would gain, the lower group perhaps most, since it is furthest from where a mixed class is usually pitched.

But something else does change, and this is the step people skip. Sorting does not only regroup students — it creates a label, and the label lands on a class that will now be taught for a year. Ask what a teacher rationally does with a class labelled low: reduce pace, simplify material, cover less ground. Each move is defensible for the class in front of them. Their cumulative effect is that the two groups are no longer covering the same curriculum, so the difference between them at the end is partly a difference in what they were offered.

Now add expectations. Teachers form beliefs about what a class can do, and those beliefs shape how much is demanded of it; students form beliefs about themselves from where they were placed. Neither effect is enormous — the classic expectancy studies are much contested and the honest summary is "real but modest" — but they push in the same direction as the curriculum difference rather than against it.

Then close the loop. Next year's placement is made from this year's attainment. A student taught less has learned less, is measured as lower, and is placed lower again. The sorting is therefore not a one-time measurement of ability but a process that partly manufactures the thing it measures — which is why the gap between tracks widens over time in the systems that track early and hard.

So what does the evidence actually say? Roughly this. The average effect of rigid between-class tracking on achievement is small — close to nothing in several syntheses, including Slavin's reviews of elementary and secondary tracking. The contested part is the distribution: several bodies of work find the top group gains modestly while the lower group loses ground, and cross-country comparisons suggest that systems selecting students very early increase inequality of outcomes without raising average performance. Meanwhile, arrangements that keep the flexibility and drop the label do better: within-class grouping for a specific subject, regrouping for one subject across classes, and above all acceleration for highly able students, which has among the more consistently positive results in education research.

Does that pattern make sense on the mechanism above? It does, and that is the useful part. Those arrangements deliver the instructional match while limiting the curriculum divergence, the permanence, and the label — they are short-term, subject-specific, and reviewed often. Which suggests the benefit and the harm come from different features of the same policy, and can partly be separated.

c

The analogy

THE ANALOGY #
THE FIGURE

Think of grouping like assigning lanes in a swimming pool. Lanes by speed genuinely help — nobody is climbing over anybody, and each swimmer can hold a sustainable pace. The trouble starts if the slow lane is also given a shorter session and a less experienced coach, and if swimmers are assigned in September and never reassessed. Then the lanes stop describing speed and start producing it.

WHERE IT BREAKS DOWN

swimmers know their times objectively and can see the fast lane's session, whereas students mostly cannot compare what they were taught with what another class was taught — so the divergence in provision is invisible to the people it affects, and shows up only in the results.

d

Clarifying the model

THE MODEL #

Three refinements, and one warning.

First, the mechanism is not "grouping is bad." It is that grouping has two consequences with opposite signs — better instructional match, and differentiated curriculum plus expectation effects — and that the balance depends on design choices, not on the sorting itself. Flexible, subject-specific, frequently reviewed grouping keeps more of the first. Rigid, whole-timetable, early tracking accumulates more of the second.

Second, a small average effect is not a small effect. An average over two groups moving in opposite directions can be near zero while both groups are substantially affected. This is why "tracking makes no difference" and "tracking matters a great deal" are both quoted from the same literature, and the second is closer to what the studies actually contain.

Third, the topic is politically loaded, and I should be plain about which parts are contested. The near-zero average is fairly robust. The claim that the top gains is reasonably supported. The size of the loss to the lower group, and how much of it is caused by grouping rather than by pre-existing differences, is genuinely argued over, because the students are not randomly assigned and placement correlates with family background — which makes the equity question, in many countries, the real dispute rather than the achievement question.

e

A picture of it

THE PICTURE #
Ability grouping
Ability grouping Start at the rounded terminal at the top and go down to the diamond, which is the only decision that matters: whether sorting changes the pace alone or changes the content and the expectations as well. The left branch is the benign version -- instructional match without divergence -- and reaches a shared gain. The right branch adds a narrowed curriculum and a label, and reaches the lopsided outcome. Both feed the small average, which is why the average conceals the story; and the arrow looping from the store back to the top is the part that makes tracking self-confirming, since next year's placement is made from a result this year's placement helped cause. {"generator":"[email protected]","source":"../Socrates/.diagram-cache/_src/ability-grouping.md","sourceIndex":1,"sourceLine":4,"sourceHash":"ff349227e441e4c86c37c06cf2aa6371e069bf807eb99977b974325f8cd9dce2","diagramType":"flowchart-v2","layoutVariant":"source","repairedDuplicateIds":[],"motion":"entrance-with-reduced-motion-fallback","presentation":"editorial","attempt":1,"viewBox":{"x":0,"y":0,"width":847,"height":1129},"qa":{"passed":true,"findings":[]}} pace only, reviewed often content and expectationstoo measured lower, placedlower again Students sorted by measuredattainment Does the sorting change what istaught? Pace matched, curriculum kept Curriculum narrowed for thelower group Expectations adjust to the label Instructional match improves Both groups can gain Top gains, lower group losesground Average effect stays small Next placement uses thisattainment
KINDSsourcedecisionprocessriskoutcomereference

How to readStart at the rounded terminal at the top and go down to the diamond, which is the only decision that matters: whether sorting changes the pace alone or changes the content and the expectations as well. The left branch is the benign version — instructional match without divergence — and reaches a shared gain. The right branch adds a narrowed curriculum and a label, and reaches the lopsided outcome. Both feed the small average, which is why the average conceals the story; and the arrow looping from the store back to the top is the part that makes tracking self-confirming, since next year's placement is made from a result this year's placement helped cause.

f

What became clearer

WHAT CLEARED #
WHAT CLEARED

Sorting students does two things at once, and they have opposite signs. Narrowing the spread genuinely helps a teacher aim, which is why acceleration and flexible within-class grouping have decent evidence behind them. But sorting also creates a label that quietly changes what is taught and what is expected, and then feeds the result back into next year's sorting — so the measurement becomes partly a cause. The average effect is small precisely because the two groups move apart, and the argument about tracking is, in the end, an argument about that distribution rather than about the mean.

g

Where to go next

ONWARD #
  • Whether peer effects — what classmates contribute directly to each other's learning — add to this or are largely captured by the teaching differences.
  • How school systems that select at age ten differ in outcomes from those that select at sixteen.
h

Key terms

TERMS #
TermWhat it means
Between-class trackingassigning students to different classes or streams by attainment, usually across most or all subjects.
Within-class groupingforming attainment groups inside a single classroom for particular work, with the same teacher and curriculum.
Accelerationmoving a highly able student to more advanced material or an older cohort rather than to a differently taught class.
Expectancy effectthe tendency for a teacher's or student's belief about likely performance to influence actual performance; real but modest, and much argued over.

Every term the collection defines is gathered in the glossary.

Nearby on the shelf

4