THIS EXPLANATION
THE ROOM
MED·31 Health & Medicine 6 MIN · 8 STATIONS

Individual treatment effect

A Socratic walk-through of the individual treatment effect — reasoned out one step at a time, not lectured.

abcdefgh
a

The question we started with

THE QUESTION #

Why can a trial prove a drug works and still never tell whether it worked for you?

A large randomised trial reports that a drug cuts the two-year event rate from twenty per cent to fifteen. You take it. Two years later you have had no event.

Did the drug work for you? The instinct is that this is a hard question — that with a bigger trial, better subgroup analysis, or a genetic test, it could be answered. I want to suggest something less comfortable: it is not a hard question of the ordinary kind. The information required to answer it is not in the data, and would not be in a trial a thousand times larger. Worth asking why.

b

Reasoning it through

REASONING #

Let us be precise about what "the drug worked for me" means. It means: what happened to me on the drug was better than what would have happened to me without it. Two outcomes for one person — the one that occurred, and the one that would have occurred under the other choice.

Now look at what you can actually observe. Exactly one of those two. You took the drug, so the untreated version of your two years does not exist anywhere to be measured. This is the fundamental problem of causal inference, named as such by Paul Holland in 1986: a causal effect is defined as a comparison of two potential outcomes for the same unit, and at most one of them is ever realised. The other is not missing in the sense of unrecorded. It never happened.

So how does a trial escape this? Ask what randomisation buys. It does not recover your missing outcome. What it does is make the control group's average outcome a fair stand-in for what the treated group's average outcome would have been — because the two groups were assembled by coin flip and so differ only by chance. That is enough to identify an average — a remarkable trick, but notice exactly what it delivers: the average of the individual effects, and nothing finer.

Push on that. Suppose the trial's five-percentage-point difference is real. What compositions of individual effects could produce it? One possibility is that the drug shifts everyone's risk slightly. Another is that it is decisive for one person in twenty and irrelevant to the rest. Another is that it strongly helps a quarter of patients and mildly harms some of the others, netting out to five points. Ask yourself which of these the trial rules out. None of them. All three yield the same two numbers, because the trial observes each patient in one arm only, and so learns the two marginal distributions of outcomes — never the joint distribution that pairs a person's treated outcome with their untreated one.

That is the crux, and it is a limit on information rather than on precision. Increasing the sample size sharpens the estimate of the average and tells you nothing whatever about the spread of individual effects, because no patient in any trial of any size contributes more than one of the two numbers you would need from them.

Would subgroups rescue it? Only partly, and in a way worth being clear about. Splitting by age or genotype gives you the average effect within that subgroup — a conditional average, still an average. You can keep conditioning on finer strata, and personalised medicine is largely the project of doing so usefully, but there is always a final step from "people like you" to "you", and that step is exactly the one the data cannot take. The honest position is that we can sometimes bound how much variation in individual effects is compatible with the observed margins, but we cannot identify it.

One genuine partial escape deserves naming. If a condition is chronic and stable, and the drug's effect is reversible with no carryover, you can treat the same person in alternating periods — an N-of-1 trial. That gets you repeated within-person comparisons, which is as close to the counterfactual as we get. It works for asthma control or chronic pain. It cannot work for a stroke, a single course of chemotherapy, or anything you only get once.

c

The analogy

THE ANALOGY #
THE FIGURE

Suppose you want to know whether your raincoat kept you dry on Tuesday. You wore it and arrived dry. To answer, you would need to know how wet you would have been on that same Tuesday walk without it — and there was only one Tuesday. You can settle the general question handsomely: send a thousand people out, coats assigned by coin flip, and measure how much drier the coated group arrives. What you can never obtain, by any amount of that effort, is the second version of your own Tuesday.

WHERE IT BREAKS DOWN

rain comes back, so you can repeat the walk next week without a coat and learn something about yourself — and that repetition is precisely what an N-of-1 trial exploits, whereas for an event you only face once there is no second Tuesday to run.

d

Clarifying the model

THE MODEL #

The misconception worth correcting gently is the idea of the "responder". People say a patient responded to a drug, meaning they improved while taking it. But improvement is not effect. Conditions fluctuate, extreme values drift back toward the middle, and people get better for reasons unrelated to what they swallowed. Distinguishing a responder from someone who was going to improve anyway requires the very comparison we have just shown is unavailable in a single person. Apparent response is a mixture, and there is no way to unmix it from one arm of one person's history.

This is also why "the trial says it works, so it will work for me" and "the trial is only an average, so it tells me nothing" are both wrong. The average is real, transferable information about a decision made under uncertainty — it is genuinely the best available basis for choosing. It simply is not a claim about your particular counterfactual, and no amount of confidence in the trial converts it into one.

Two limits on my own account. Effect heterogeneity is not entirely beyond reach: a crossover design or strong modelling can reveal something of its structure, always at the price of assumptions the data cannot check. And the framework treats the effect as well defined for each person, which presumes that treatment means the same thing for everyone.

e

A picture of it

THE PICTURE #
Individual treatment effect
Individual treatment effect the left node is one arm of a trial; the four bands are the composition a clinician actually wants. The bands are illustrative rather than measured, and that is the point -- the trial reports only the net difference between arms, here twenty helped minus five harmed, and every set of bands with the same net produces an identical trial result. The picture shows the thing that is invisible, not the thing that was found. {"generator":"[email protected]","source":"../Socrates/.diagram-cache/_src/individual-treatment-effect.md","sourceIndex":1,"sourceLine":4,"sourceHash":"bbac6a98ebd4c288b3f891e2d6918870774988ef8ebd96f93668b3e7ae965db9","diagramType":"sankey","layoutVariant":"source","repairedDuplicateIds":[],"motion":"entrance-with-reduced-motion-fallback","presentation":"editorial","attempt":1,"viewBox":{"x":0,"y":0,"width":720,"height":542},"qa":{"passed":true,"findings":[]}} Treatedarm · 100 Helpedbydrug · 20 Betteranyway · 60 Unaffected · 15 Harmedbydrug · 5

How to readthe left node is one arm of a trial; the four bands are the composition a clinician actually wants. The bands are illustrative rather than measured, and that is the point — the trial reports only the net difference between arms, here twenty helped minus five harmed, and every set of bands with the same net produces an identical trial result. The picture shows the thing that is invisible, not the thing that was found.

f

What became clearer

WHAT CLEARED #
WHAT CLEARED

A randomised trial answers a question about a population, and it answers it well. The question about you is a different question, and it asks for a comparison to a version of your life that did not occur — so the gap between the two is not a shortfall in evidence to be closed by more evidence, but a boundary on what any observation of a single person can contain.

g

Where to go next

ONWARD #
  • The Frechet-Hoeffding bounds, which say how much can be inferred about individual effects from marginals alone.
  • How N-of-1 trials are designed, and which conditions genuinely qualify for them.
h

Key terms

TERMS #
TermWhat it means
Potential outcomesthe pair of results a person would have under treatment and under no treatment; the difference between them is the individual treatment effect.
Fundamental problem of causal inferencethe fact that only one potential outcome per person is ever observed, so an individual effect is never directly measurable.
Average treatment effectthe mean of the individual effects; what randomisation identifies.
Effect heterogeneityvariation in the individual treatment effect across people, compatible with any given average.
N-of-1 triala randomised, often blinded sequence of treatment and control periods within a single patient.

Every term the collection defines is gathered in the glossary.

Nearby on the shelf

4