THIS EXPLANATION
THE ROOM
MED·35 Health & Medicine 6 MIN · 8 STATIONS

Randomized trials

A Socratic walk-through of randomized trials — reasoned out one step at a time, not lectured.

abcdefgh
a

The question we started with

THE QUESTION #

Why does deciding treatment by coin toss give more trustworthy answers than choosing carefully?

An experienced physician knows her patients: who is frail, who will tolerate the harder drug, who is likely to do badly whatever happens. We are asking her to set all of that aside and let a random number decide who gets which treatment.

That sounds like deliberately throwing away information — and it is. So the question is not "why is randomising harmless?" but the sharper one: what exactly does the coin buy that is worth the knowledge we discard, and what does it conspicuously fail to buy?

b

Reasoning it through

REASONING #

Begin with what a comparison requires. To read a difference in outcomes as an effect of treatment, the two groups must differ in the treatment and in nothing else that bears on the outcome.

Now let the clinician assign. She is good at her job, so her assignment is a function of prognosis: the demanding surgery goes to those fit enough to survive it. Treatment and outcome now share a common cause, and if the surgical group does better we cannot separate the surgery from the fitness that earned a place in it. Notice that this contamination is worst precisely when her judgement is best, because a stronger predictor of outcome is being fed into the assignment.

Could we not measure the confounders and adjust? Partly. Age, stage, kidney function, comorbidity — all can be regressed out. But adjustment reaches exactly as far as measurement does, and the strongest prognostic variable in the room is often the clinician's unrecorded impression of how the patient looks. It went into the assignment and never into the dataset.

So what does a coin do? Here is the step that is most often stated wrongly. It is tempting to say randomisation balances the groups. It does not, and it cannot. In any single trial the two arms will differ on countless variables, some of them substantially, and with small numbers this is not a rare accident but the normal case.

What randomisation actually does is make allocation independent of the patient — of everything about them, recorded or not, known or not yet discovered. That is a claim about the mechanism, not about the resulting groups. Its consequence is that any imbalance which does arise was produced by the coin rather than by prognosis: it is chance, not selection. The guarantee is in expectation, over the ensemble of trials the procedure could have generated, and never for the trial in front of you.

That distinction earns its keep immediately, because it hands over the second thing the coin buys. If the imbalances come from a known chance mechanism, then the size of imbalance chance alone could produce is calculable. That is where a p-value or a confidence interval comes from — the randomisation supplies the probability model the test is computed under. Without a known allocation mechanism, the sampling distribution is a story we tell.

There is a third purchase, easy to miss. A random sequence is also an unpredictable one — but only if nobody can see ahead. If the recruiter can work out the next allocation, they can hold a frail patient back for the arm they would rather protect, and prognosis re-enters an impeccably randomised trial. Hence concealing the upcoming allocation is a requirement distinct from randomising at all. Comparisons across many published trials have found that those with inadequately concealed allocation report larger treatment effects — on the order of a third larger, a recalled magnitude and much the softest number here.

c

The analogy

THE ANALOGY #
THE FIGURE

Picture a farmer testing a fertiliser across a field. She spreads it where the crop looks poorest, which is exactly what a sensible farmer does. At harvest the treated strips yield less, and the comparison is unreadable — no care in weighing the grain can repair it. Toss a coin per strip instead and the strips are still not identical; one arm may get more of the wet corner. But the assignment now knows nothing about the soil, and the difference the coin alone could have produced is a quantity she can work out.

WHERE IT BREAKS DOWN

fields do not drop out of the trial, ask which strip they were in, or receive extra water because the farmer expects them to fail — and nearly every way a randomised trial actually goes wrong happens after the sowing, exactly where this analogy falls silent.

d

Clarifying the model

THE MODEL #

Take that last point seriously, because it bounds the whole method. Randomisation protects one instant, the moment of allocation, and says nothing about what follows. Differential dropout, outcome assessors who can see which arm a patient is in, extra care given to the group a team believes in, and stopping early on a favourable interim look all reintroduce bias into a perfectly randomised trial. Blinding, complete follow-up and a pre-specified analysis are separate defences, not consequences of the coin.

Two further limits. Randomisation buys internal validity only: the answer applies to the sort of patient actually enrolled, and trial populations are usually younger and more adherent than the clinic's. And it certifies nothing about whether the outcome measured matters — a trial can be immaculately randomised and report a surrogate nobody feels.

One correction to what the model is often taken to imply: because imbalance in a randomised trial is chance by construction, significance-testing the baseline table is incoherent — you would be testing a hypothesis you already know to be true.

Now the honest attempt to break my own claim. I have argued that the trustworthiness comes from allocation being independent of prognosis. If that is right, careful non-randomised comparisons should sometimes be badly wrong in ways their authors could not detect; if instead they routinely agreed with subsequent trials, the whole argument would be an elegant irrelevance. The record supplies the test rather than my reasoning: observational cohorts indicating that hormone replacement protected against coronary disease were followed by a large randomised trial that did not find that benefit, and the discrepancy was traced to who chose to take it. A run of the opposite result would have sunk the case.

e

A picture of it

THE PICTURE #
Randomized trials
Randomized trials This repurposes a version-history graph: the trunk is one group of patients over time, the branch point is the randomisation, and the two lines are the arms. Everything before the split is shared history, which is the whole source of the method's authority -- the arms are identical up to that commit because nothing about the patient influenced where they went. Read the branches downstream as the part the coin does not protect, and the join at the end as the analysis comparing the arms, not the patients rejoining. {"generator":"[email protected]","source":"../Socrates/.diagram-cache/_src/randomized-trials.md","sourceIndex":1,"sourceLine":4,"sourceHash":"443ad3e575cb87dd45dda00deb8bd5b649e83a6d950609b35de7f9851a2b9196","diagramType":"gitGraph","layoutVariant":"source","repairedDuplicateIds":[],"motion":"entrance-with-reduced-motion-fallback","presentation":"editorial","attempt":1,"viewBox":{"x":0,"y":0,"width":725,"height":378},"qa":{"passed":true,"findings":[]}} main treated control eligible baseline coin toss drug dropouts A outcome A placebo dropouts B outcome B

How to readThis repurposes a version-history graph: the trunk is one group of patients over time, the branch point is the randomisation, and the two lines are the arms. Everything before the split is shared history, which is the whole source of the method's authority — the arms are identical up to that commit because nothing about the patient influenced where they went. Read the branches downstream as the part the coin does not protect, and the join at the end as the analysis comparing the arms, not the patients rejoining.

f

What became clearer

WHAT CLEARED #
WHAT CLEARED

The coin is not a way of making two groups the same. It is a way of guaranteeing that no fact about a patient — including the ones nobody wrote down, and the ones nobody has thought of yet — had any hand in deciding their treatment. What we give up is the clinician's judgement in the assignment; what we get back is an assignment that carries none of her information into the comparison, plus a known chance model to measure the leftover noise against.

Read that way, the method is narrower than its reputation. It secures a single moment extremely well and leaves everything around it — who was enrolled, who was lost, what was measured, when it was stopped — to be argued about.

g

Where to go next

ONWARD #
  • How to reason about a treatment that cannot ethically be randomised, and what instrumental variables attempt there.
  • Why pre-registration and pre-specified analysis matter as much as the allocation itself.
h

Key terms

TERMS #
TermWhat it means
Confoundinga variable influencing both who gets the treatment and how they fare, making the raw comparison unreadable.
Allocation concealmentkeeping the upcoming assignment hidden from whoever enrols patients, so the sequence cannot be gamed.
Blindingkeeping patients, clinicians or assessors unaware of the assigned arm, protecting the period after allocation.
Intention to treatanalysing every patient in the arm they were allocated to, preserving the randomisation despite non-adherence.

Every term the collection defines is gathered in the glossary.

Nearby on the shelf

4