Rubber-stamped review
A Socratic walk-through of rubber-stamped review — reasoned out one step at a time, not lectured.
The question we started with
THE QUESTION #Why does a human reviewer who approves almost everything provide almost no safety?
The reassurance is offered in almost every design review: "there is a human in the loop." Someone sees each action before it takes effect. It sounds decisive, and in the incident report six months later it will turn out that the human approved the harmful action in about two seconds, along with the four hundred harmless ones either side of it.
So let us not ask whether human review is good. Let us ask a narrower question: what would have to be true of a reviewer for their approval to carry information? Because if approval is unconditional, it tells you nothing about the thing approved — and a signal that is always the same is not a signal.
Reasoning it through
REASONING #Start with arithmetic, because it settles more of this than psychology does.
Suppose ninety-eight per cent of the items reaching the reviewer are fine. What accuracy does a reviewer achieve by approving everything without looking? Ninety-eight per cent. They will appear excellent on any measure that counts agreement with the eventual outcome, and their manager's dashboard will show a near-perfect record. Does that dashboard distinguish them from a reviewer who is genuinely catching things?
It does not, and that is the first trap: the metric that is easy to collect is exactly the one that cannot see the failure. The only informative number is the conditional one — of the bad items that reached this reviewer, what fraction did they stop? And that number is expensive to compute, because you have to know which items were bad independently of the review.
Now the psychology, which explains why the arithmetic goes unnoticed. Two effects are well documented. Automation bias is the tendency to accept an automated system's recommendation and to search less hard for disconfirming evidence once a machine has proposed an answer — studied since the 1990s in aviation and clinical decision support, where it produces both errors of commission, following a wrong recommendation, and errors of omission, failing to notice something the machine did not flag. And the vigilance decrement: sustained attention for rare signals decays within tens of minutes, which is why watch-keeping tasks with very low event rates are among the hardest jobs to do well.
Put those together with the base rate. The reviewer's honest experience is that the system is nearly always right. Approving is rewarded with throughput; rejecting costs time, invites an argument, and is usually wrong. What behaviour would you predict emerges over a few thousand items?
There is a third layer, and it is the one that makes this a question about incentives rather than about attention. What is the review for, institutionally? Often, in truth, it is there so that a person's name is attached to the decision. That function is fully served by the approval whether or not anyone looked. A control whose organisational purpose is discharged by the signature alone will decay to the signature alone — not because reviewers are lazy, but because nothing in the system distinguishes the two states.
The analogy
THE ANALOGY #Think of an airport X-ray operator. The job is to spot a weapon in a bag, and in ordinary operation there are none for weeks. Left alone, attention collapses — not through carelessness but because nothing in the task suggests looking is ever worthwhile. Which is why real screening systems inject fictional threat images into the stream: the operator now meets something every hour or so, their catch rate becomes measurable, and looking gains feedback rather than remaining an act of faith.
an injected X-ray image is indistinguishable from the real thing and the operator's task is unchanged, whereas synthetic bad cases in a model-review queue are much harder to make realistic — and a reviewer who learns to recognise the fakes has learned the wrong skill entirely.
Clarifying the model
THE MODEL #The reframing that helps is to stop treating review as a gate and start treating it as a measurement. A gate is binary and unfalsifiable; a measurement has a detection rate, a false-alarm rate, and a cost per item, and those can be designed.
Once you look at it that way, the fixes stop being exhortations to be careful.
Change what is reviewed. Reviewing everything guarantees the base rate is terrible and the attention budget is spread evenly over items that do not need it. Reviewing a triaged subset — the uncertain, the unusual, the high-consequence — raises the density of genuine decisions per minute of human time, which is the quantity that determines whether the human is doing anything.
Change what is measured. Track catch rate, not throughput — which requires known-bad items in the stream, or an independent audit of sampled approvals. Without one of those, you cannot tell a functioning control from a decorative one until the incident.
Change the action. Approving is a single click and rejecting is an essay — so make the effort symmetric, or require the approver to record the specific thing they checked. A review that asks "does this cite a real source?" produces attention; a review that asks "OK?" produces clicks.
Change the consequences. If the reviewer is measured on queue length and blamed for delays but never rewarded for a catch, no amount of training will overcome the incentives. This is where the fairness question sits too: placing a human at the point of approval concentrates blame on the person least able to see the whole system — what Madeleine Elish named the moral crumple zone in 2019.
One misconception worth correcting directly: the problem is not that humans are bad reviewers. Given a hard case, adequate time, the relevant evidence surfaced, and a reason to believe their judgement matters, humans are excellent. Every one of those four conditions is a design decision, and rubber-stamping is what happens when all four are left to chance.
A picture of it
THE PICTURE #How to readthe widths are counts from an illustrative thousand items in which twenty are genuinely harmful, and the reviewer inspects six per cent of what arrives. Follow the two streams out of the queue: the thick band flows onward almost untouched, carrying nineteen of the twenty harmful items with it, while the thin examined band stops one. The lesson is in the proportions — the catch rate ends up equal to the inspection rate, so the control's effectiveness is set by how much attention each item receives, not by the fact that every item passed a human. Note too how large the "fine anyway" flows are: that volume is what makes an approve-everything habit look, on any simple metric, like excellent performance.
What became clearer
WHAT CLEARED #Approval carries information only when refusal was a live possibility, and a reviewer facing a stream that is nearly always fine, measured on throughput, and clicking a single button has no live possibility of refusal left. So "a human is in the loop" is not a safety property at all — the safety property is the reviewer's measured catch rate on cases that genuinely mattered, and if nobody is measuring that, nobody knows whether the loop is closed.
Where to go next
ONWARD #- Signal detection theory: separating a reviewer's discrimination ability from their decision threshold.
- Triage and selective review — deciding which items deserve human attention when the budget is a fixed number of minutes.
- Automation bias in clinical decision support, where the pattern has real patient outcomes attached.
Key terms
TERMS #| Term | What it means |
|---|---|
| Automation bias | the tendency to over-trust an automated recommendation and to reduce independent checking once it is offered. |
| Vigilance decrement | the well-documented decay of sustained attention on tasks where the signal to be detected is rare. |
| Base rate | the underlying proportion of bad items, which sets how good a do-nothing reviewer looks on naive metrics. |
| Catch rate | the fraction of genuinely bad items a reviewer stops; the only measure that separates a working control from a decorative one. |
| Moral crumple zone | Madeleine Elish's term for a human placed where they will absorb blame for a failure they had little practical capacity to prevent. |
Every term the collection defines is gathered in the glossary.