THIS EXPLANATION
THE ROOM
EDU·37 Education & Learning 6 MIN · 8 STATIONS

Tutoring versus class size

A Socratic walk-through of tutoring versus class size — reasoned out one step at a time, not lectured.

abcdefgh
a

The question we started with

THE QUESTION #

Why does one-to-one tutoring transform results when cutting a class from thirty to twenty-five barely helps?

Everyone believes both halves of this: a private tutor works wonders, and smaller classes are better. The second belief is the expensive one — cutting classes by a fifth means a fifth more teachers and the rooms to put them in — and when it is done and measured, the gains are disappointing to the point of embarrassment. Yet tutoring, the same idea taken to its limit, plainly does something. If the ratio is what matters, why is the curve so flat until the very end of it?

b

Reasoning it through

REASONING #

Start with the claim that anchors the debate. In 1984 Benjamin Bloom reported that students taught one-to-one, with mastery learning, performed about two standard deviations above conventionally taught students — roughly moving an average student to the 98th percentile. That number, "2 sigma", is quoted constantly. It has not replicated at that magnitude. Later syntheses of human tutoring put the effect nearer 0.8 standard deviations, and the largest recent evidence base, a pooled analysis of dozens of randomised trials, lands near 0.4. Those are large effects by education standards, but they are not 2 sigma, and treating the original figure as a benchmark has distorted much thinking.

Now the other side. The Tennessee STAR experiment of the late 1980s randomly assigned young children to classes of roughly 13 to 17 or 22 to 25, and found gains of about 0.15 to 0.2 standard deviations, larger for disadvantaged pupils and concentrated in the earliest grades — a real effect, from a cut of about a third. When California cut class sizes statewide a decade later the measured benefit was slight, partly because staffing thousands of new classrooms at once meant hiring less experienced teachers.

Put the two together and a shape appears: something is very strong at one-to-one, weak at a third-sized cut in early grades, and undetectable at thirty-to-twenty-five. What kind of mechanism produces that?

Not a ratio. A ratio is a smooth quantity, and a smooth quantity should give you a fifth of the benefit for a fifth of the cut. So ask what a tutor has that a teacher of twenty-five lacks, and check whether the teacher of thirty acquires any of it by dropping to twenty-five.

Watch a tutor for ten minutes. The student attempts something; the tutor sees, in seconds, exactly where it went wrong — not that it was wrong, but which step, which misconception — and responds before the error is rehearsed. The next problem is chosen for this student's gap. Nothing is aimed at a median that does not exist. And silence is immediately interrogated, so nobody can sit quietly and be counted as following.

Every one of those is a property of the feedback loop, not of the headcount: how fast the teacher learns each student's state, how specific that information is, and how quickly instruction adjusts to it. In a class of thirty the information is coarse, delayed, and sampled from the few who answer — the loop closes at the pace of marked homework, days later, in aggregate. In a class of twenty-five it closes the same way. The mode of instruction is unchanged, so the loop is unchanged, so the outcome barely moves.

That also explains STAR. A class of fifteen five-year-olds is not merely 40 per cent smaller; it is small enough to change what the teacher does — more individual attention, less time on the behaviour management that dominates the youngest classrooms. Where the cut changes practice, results move. Where it only changes a number, they do not.

c

The analogy

THE ANALOGY #
THE FIGURE

Think of a thermostat. A room controlled by a sensor that reports the average temperature once a day will be badly regulated no matter how powerful the heater. Halve the reporting interval to twelve hours and it is still badly regulated. What transforms the control is a sensor that reads continuously and a valve that responds at once — and until you get there, improvements in the size of the room or the strength of the heater hardly show up in how steady the temperature is.

WHERE IT BREAKS DOWN

a thermostat has one variable to regulate and one target; a class has thirty learners with different states and no single setpoint, which is exactly why the sampling problem is severe. And a student, unlike a room, is not passive — part of what tutoring changes is that being observed makes disengagement impossible, which has no thermostatic counterpart.

d

Clarifying the model

THE MODEL #

The claim is not that class size is irrelevant. It is that class size acts on outcomes only through what it lets a teacher do, and that relationship is lumpy rather than proportional — small cuts change nothing about the method, so they buy little; large cuts in early grades change the method, so they buy something real but modest.

Several honest qualifications. Effect sizes from different studies are not strictly comparable: they rest on different tests, ages and comparison conditions, and the small trials producing the largest numbers are exactly those most inflated by publication bias. Tutoring effects vary enormously with who does it — trained teachers outperform paraprofessionals, who outperform volunteers and parents — and with dosage, subject and grade. And tutoring at scale is not the tutoring in these trials: the programmes that worked ran in school hours with trained, supervised tutors, and staffing that nationally would meet the same dilution problem that blunted the Californian reform.

The practical inference is not "hire tutors instead of teachers", but that an intervention should be judged by whether it tightens the loop — frequent low-stakes checks and responsive re-teaching move information faster without a tutor for every child.

e

A picture of it

THE PICTURE #
Tutoring versus class size
Tutoring versus class size Each bar is an estimated effect in standard deviations, so 0.2 is a modest but real shift and 1.0 would be enormous. Read left to right as a descent from the famous claim to the measured reality: Bloom's 1984 two-sigma figure, then human tutoring as later syntheses put it, then the pooled average across many randomised trials, then the STAR cut from around 23 pupils to around 15 in early grades. The final bar sits near the floor deliberately -- a cut from thirty to twenty-five has no reliably measurable effect, and its height stands for "too small to distinguish from zero" rather than a measured value. {"generator":"[email protected]","source":"../Socrates/.diagram-cache/_src/tutoring-versus-class-size.md","sourceIndex":1,"sourceLine":4,"sourceHash":"9c5af91689e3a56db58081fac79351e5fd5b58fd67bc79c5fa60d84e4689b793","diagramType":"xychart","layoutVariant":"source","repairedDuplicateIds":[],"motion":"entrance-with-reduced-motion-fallback","presentation":"editorial","attempt":1,"viewBox":{"x":0,"y":0,"width":790,"height":636},"qa":{"passed":true,"findings":[]}} Bloom claim Best tutoring Tutoring mean STAR small 30 to 25 2.2 2 1.8 1.6 1.4 1.2 1 0.8 0.6 0.4 0.2 0 Effect size

How to readEach bar is an estimated effect in standard deviations, so 0.2 is a modest but real shift and 1.0 would be enormous. Read left to right as a descent from the famous claim to the measured reality: Bloom's 1984 two-sigma figure, then human tutoring as later syntheses put it, then the pooled average across many randomised trials, then the STAR cut from around 23 pupils to around 15 in early grades. The final bar sits near the floor deliberately — a cut from thirty to twenty-five has no reliably measurable effect, and its height stands for "too small to distinguish from zero" rather than a measured value.

f

What became clearer

WHAT CLEARED #
WHAT CLEARED

Tutoring and class-size reduction are not the same lever at different settings. Tutoring works because it replaces a slow, coarse, sampled feedback loop with a fast and specific one; class-size reduction only helps when it is large enough to force that same change in practice, which is why it shows up in early grades at a third off and vanishes at a fifth off. And the number everyone repeats, two sigma, is the least reliable figure in the whole discussion.

g

Where to go next

ONWARD #
  • Why formative assessment and mastery learning show sizeable effects without any change in staffing ratio.
  • Whether computer tutoring systems close the gap to human tutors, and what they still cannot read in a student.
h

Key terms

TERMS #
TermWhat it means
Effect size (standard deviation)a measure of an intervention's impact expressed in units of the spread of student outcomes, allowing comparison across different tests.
Bloom's 2 sigma problemthe 1984 finding, since not replicated at that size, that tutored students outperformed conventionally taught ones by about two standard deviations.
Project STARthe randomised Tennessee experiment of the late 1980s assigning early-grade pupils to small or regular classes.
Formative assessmentchecking understanding during instruction in order to adjust it, rather than to grade it afterwards.

Every term the collection defines is gathered in the glossary.

Nearby on the shelf

4