THIS EXPLANATION
THE ROOM
EDU·03 Education & Learning 6 MIN · 8 STATIONS

Ceiling on admissions prediction

A Socratic walk-through of the ceiling on admissions prediction — reasoned out one step at a time, not lectured.

abcdefgh
a

The question we started with

THE QUESTION #

Why can no admissions test tell which applicant will actually flourish once admitted?

Every few years someone proposes a better admissions instrument. A longer test, a structured interview, a personality inventory, a machine-learning model over the whole application. Each is an improvement on the last in some respect, and each disappoints in the same way: it sorts applicants a little better than chance and nowhere near well enough to identify the ones who will thrive.

The natural reading is that we have not yet found the right measure. But notice how odd that reading becomes after a century of trying, with enormous incentives and very large datasets. It is worth entertaining the alternative: that the ceiling is not in the instruments but in what is available to be measured at the moment of the decision. If that is true, then the question is not which test to build but what a test could possibly be reading — and that is a question we can actually reason about.

b

Reasoning it through

REASONING #

Start at the far end, with the thing we want to predict. What is it? "Flourishing" is not a quantity anyone has recorded. Systems that study this substitute something they do have — first-year grade average, mostly, sometimes graduation, occasionally earnings a decade later. Already something has gone wrong that has nothing to do with the test. Grades are themselves noisy and depend on which courses a student picked and who marked them, so even a perfect predictor of the underlying quality would correlate imperfectly with the recorded outcome. An unreliable target puts a hard cap on any correlation with it, no matter how good the instrument.

Second, ask what data the correlation is even computed on. Only admitted students have outcomes. The whole lower half of the applicant pool — the people the test says would struggle — is never observed, so the visible spread of ability is compressed. Compress the spread of a predictor and its measured correlation drops mechanically, which is why validity coefficients computed inside a selective institution understate what the test does across the full range. Psychometricians call this range restriction and correct for it, but the correction is a model, not an observation.

Now the step that matters most, and the one that changes the character of the problem. Ask when the outcome is caused. A student flourishes or founders partly because of what they brought, and substantially because of what happened afterwards: whether the money held out, whether they were taught by someone who lit them up, whether they fell ill, fell in love, chose the wrong course, found a community or did not. Those causes had not occurred when the application was read. They were not weakly recorded or poorly measured — they did not exist. No instrument, however sophisticated, can extract information from a record that does not contain it, and no volume of training data changes that, because the missing variables are missing for every applicant in the dataset too.

That is the ceiling, and it is worth stating plainly because it is easy to slide past: the limit is informational, not technological. A better test can recover more of what is present in the applicant. It can recover none of what is not yet true about the world.

Two smaller forces push the ceiling down further. One is that the measure is used for selection, so applicants and the industry around them optimise against it; coaching moves scores without moving the thing scores were supposed to indicate, and the more consequential a measure becomes the faster this erodes it. The other is that admission is not a prediction so much as an intervention — it changes the student's environment, peers and expectations, so the thing being forecast is partly created by the forecast.

Where does that leave the numbers? Roughly here, and I would hold the figures loosely because they vary by institution and cohort: the best combinations, typically prior grades together with a test, correlate with first-year grades somewhere in the region of one-half after correction, which is a meaningful signal explaining perhaps a quarter of the variation. Against longer-run outcomes the relationship weakens further. That is a real, useful, and badly oversold amount of information.

c

The analogy

THE ANALOGY #
THE FIGURE

Predicting who will flourish is like judging from a photograph of a seedling how tall the tree will be in thirty years. You can read a great deal from the picture — the species, the vigour, the health of the leaves — and it genuinely narrows the range. What you cannot read is the drought in year nine, the neighbour who builds a wall, or the storm, because at the moment of the photograph none of them has happened.

WHERE IT BREAKS DOWN

a seedling does not know it is being photographed, whereas applicants can and do cultivate exactly the features the photograph rewards, so an admissions measure decays with use in a way that a botanical one does not.

d

Clarifying the model

THE MODEL #

The refinement connecting the steps is that four different limits are being confused when people say a test "does not work". A noisy outcome, a truncated sample, causes that arrive after the decision, and a measure corrupted by being used are separate problems with separate fixes — and only the first two can be attacked by better measurement at all.

The misconception worth dismantling is the inference from a modest correlation to the conclusion that the test is worthless. A predictor that explains a quarter of the variation is close to useless for any individual and quite informative across a large cohort. That gap is the source of most of the argument: institutions defend the measure with the population claim and applicants attack it with the individual one, and both are correct about their own question. Nothing about the ceiling says the information is zero; it says the residual is large and irreducible.

It also follows that the useful design response is not a sharper filter but a system that tolerates being wrong — broad first-year courses, transfer routes, and second chances. If a quarter of the variation is all that can be seen in advance, the remaining three-quarters has to be handled after admission or not at all.

e

A picture of it

THE PICTURE #
Ceiling on admissions prediction
Ceiling on admissions prediction each branch is a separate reason the ceiling exists, not a stage in a process. Better instruments can push on the first two; the third is closed to any instrument, and the fourth gets worse the more the instrument is trusted. {"generator":"[email protected]","source":"../Socrates/.diagram-cache/_src/ceiling-on-admissions-prediction.md","sourceIndex":1,"sourceLine":4,"sourceHash":"db83f862813f88d6f1922c5345d4982de95724f3dfea36b6034632b131217fb7","diagramType":"mindmap","layoutVariant":"source","repairedDuplicateIds":[{"original":"mermaid-db83f862813f88d6-0-node_1","replacement":"mermaid-db83f862813f88d6-0-node_1--duplicate-2"},{"original":"mermaid-db83f862813f88d6-0-node_2","replacement":"mermaid-db83f862813f88d6-0-node_2--duplicate-2"},{"original":"mermaid-db83f862813f88d6-0-node_3","replacement":"mermaid-db83f862813f88d6-0-node_3--duplicate-2"},{"original":"mermaid-db83f862813f88d6-0-node_4","replacement":"mermaid-db83f862813f88d6-0-node_4--duplicate-2"},{"original":"mermaid-db83f862813f88d6-0-node_5","replacement":"mermaid-db83f862813f88d6-0-node_5--duplicate-2"},{"original":"mermaid-db83f862813f88d6-0-node_6","replacement":"mermaid-db83f862813f88d6-0-node_6--duplicate-2"},{"original":"mermaid-db83f862813f88d6-0-node_7","replacement":"mermaid-db83f862813f88d6-0-node_7--duplicate-2"},{"original":"mermaid-db83f862813f88d6-0-node_8","replacement":"mermaid-db83f862813f88d6-0-node_8--duplicate-2"},{"original":"mermaid-db83f862813f88d6-0-gradient","replacement":"mermaid-db83f862813f88d6-0-gradient--duplicate-2"}],"motion":"entrance-with-reduced-motion-fallback","presentation":"editorial","attempt":1,"viewBox":{"x":0,"y":0,"width":1198,"height":620},"qa":{"passed":true,"findings":[]}} Prediction ceiling Noisy outcome Grades measure loosely Range restriction Only admits observed Later causes Events after entry Measure decays Coaching the score

How to readeach branch is a separate reason the ceiling exists, not a stage in a process. Better instruments can push on the first two; the third is closed to any instrument, and the fourth gets worse the more the instrument is trusted.

f

What became clearer

WHAT CLEARED #
WHAT CLEARED

The failure is not of testing but of availability. Much of what decides whether someone flourishes has not happened yet when the decision is made, so it is absent from every application ever written and from every dataset built from them. A test can read the applicant well and still be blind to the years that do most of the work — which makes the honest goal not a better filter but a decision that is cheap to be wrong about.

g

Where to go next

ONWARD #
  • Whether structured interviews add anything beyond prior grades, or mostly repeat them.
  • How medical and pilot selection handle the same ceiling when the cost of error is far higher.
  • What a system optimised for correcting admission mistakes rather than avoiding them would look like.
h

Key terms

TERMS #
TermWhat it means
Predictive validityhow well scores taken before an event correlate with the outcome they are meant to forecast.
Range restrictionthe mechanical shrinking of a measured correlation when only part of the applicant range is observed.
Criterion problemthe difficulty that the outcome being predicted is itself vaguely defined and unreliably recorded.

Every term the collection defines is gathered in the glossary.

Nearby on the shelf

4