Anonymous telescope-time proposals
A Socratic walk-through of anonymous telescope-time proposals — reasoned out one step at a time, not lectured.
The question we started with
THE QUESTION #Why did hiding the applicants' names change which astronomers were awarded telescope time?
Time on the Hubble Space Telescope is one of the most contested resources in science: in a typical cycle, roughly one proposal in five or six succeeds. For most of the telescope's history, panels read proposals with the applicants' names attached. Across many of those cycles, proposals led by men were selected at a slightly higher rate than proposals led by women — a gap of a couple of percentage points, small in any single cycle, but persistent enough across cycles to be hard to dismiss as noise.
In 2018, for Cycle 26, the Space Telescope Science Institute switched to dual-anonymous review: the reviewers did not know who had written the proposal, and the writers did not know who would read it. In that first anonymous cycle the gap closed and slightly reversed, with women-led proposals succeeding at a marginally higher rate than men-led ones. NASA extended the approach to other programmes, and other facilities have followed.
Now, the tempting reading is that a panel of prejudiced astronomers was caught out. I want to resist that reading, because it makes the result smaller than it is. Suppose everyone on those panels was acting in good faith. Would you still expect the change?
Reasoning it through
REASONING #Ask first what a reviewer is actually being asked to do. Not simply "is this good science" — also "will this team deliver it?" Feasibility is a legitimate and explicitly stated criterion, because telescope time spent on a programme that cannot be executed is time destroyed. So a reviewer is being asked, in part, to forecast.
Now, forecasting is expensive. A panellist may have thirty or fifty proposals to assess, each a dense technical document in a subfield that is not quite their own. What would you reach for? The cheapest available predictor of whether a team can execute is the team's record — and under named review, that predictor is sitting right there on the front page, free, before you have read a word of the science.
So notice what we have established without invoking any hostility at all. The rational strategy for a time-pressed reviewer, under the criteria they were given, is to lean on the name. The name is not noise; it genuinely carries information. It is simply that it carries the wrong information more often than anyone would like, because a track record is itself the accumulated output of earlier allocations, earlier institutions, earlier access — and it correlates with things nobody intended to select on.
Now ask the sharper question: when you remove the name, what actually changes? Not the reviewer's beliefs. What changes is the evidence available to act on. If the cheap proxy is gone, the only thing left to reward is the thing on the page — the justification, the method, the observing strategy. The reviewer's optimal strategy shifts because their information set shifted, and their behaviour follows.
And this reaches back up the chain, which is the part that gets least attention. Under named review, a proposal writer's effort is well spent on establishing standing: cite yourself, name your collaborators, signal the pedigree of the team. Under anonymous review that effort is wasted — the rules require the science case to be written without self-identifying references — so the same effort has to be redirected into the argument itself. The incentive that changed was not only the reviewer's, but the applicant's.
There is also a quieter finding that supports this reading better than the headline one: first-time and less established proposers fared better under anonymity too. That is what you would expect if the mechanism is "a reputational proxy was removed", and not what you would expect if the mechanism were about a single demographic variable alone.
I should be careful about how much weight the numbers bear. The success-rate gaps involved are of a few percentage points and the cycle-to-cycle scatter is real, so no single cycle proves anything; the case rests on a long pattern before the change and a break at it. And genuine anonymity is imperfect — more on that shortly.
The analogy
THE ANALOGY #Think of an orchestra audition held behind a screen. The panel is not being told to ignore reputation; it is being made unable to consult it, so the only thing capable of earning a place is the sound coming through the curtain. The players know this too, so preparation shifts away from cultivating the panel and toward the playing itself. Blind auditions became widespread in American orchestras from the 1970s and 1980s, and are often credited with part of the rise in women hired — though how large that effect was is genuinely argued about among economists.
A screen conceals a performer completely, whereas a proposal cannot conceal its subject, its instrument configuration, or the very particular expertise it displays — so in a small subfield an experienced reviewer can often guess the author anyway, and the anonymity is a reduction in the strength of the signal rather than its removal.
Clarifying the model
THE MODEL #The model to correct is "bias was blocked". What was actually blocked was a proxy. Reviewers were using a legitimate and informative shortcut for a criterion they were told to assess, and that shortcut carried in correlations with gender, seniority and institution that nobody had chosen to select on. Anonymity did not change anyone's values; it changed what the cheapest strategy for a reviewer was, and behaviour followed the incentive.
That framing also explains why the change was not free. Feasibility genuinely does need assessing, so anonymous systems have to reintroduce it elsewhere — separately, after the science ranking, and with the identities revealed only at that point. Removing an informative signal costs you something real; the argument for doing it is that what the signal cost was larger.
A picture of it
THE PICTURE #How to readStart at the input node at the top and take the single decision. The right-hand branch is the named regime: the reviewer reaches for the cheapest available predictor, which feeds both the allocation and the hazard node hanging beneath it. The left-hand branch is the anonymous regime, where that shortcut is unavailable and the argument on the page is the only thing left to reward. Both branches end at the same store of allocated time — the change is in what earned it, not in who decided.
What became clearer
WHAT CLEARED #Anonymising the proposals did not make reviewers fairer people. It removed a cheap, informative, contaminated shortcut, and in doing so changed what the best strategy was on both sides of the process — what a reviewer could economically rely on, and what a writer could profitably spend effort on. The shift in who won time is what a change in incentives looks like when nobody's intentions have changed at all.
Where to go next
ONWARD #- How feasibility is assessed after an anonymous science ranking, and whether reintroducing identity late lets the old effect back in.
- Whether anonymity holds up in very small subfields where a proposal's topic identifies its author.
Key terms
TERMS #| Term | What it means |
|---|---|
| Dual-anonymous review | a scheme in which neither reviewer nor applicant knows the other's identity. |
| Oversubscription rate | the ratio of requested to available telescope time; for Hubble, roughly five or six to one. |
| Proxy | an easily observed signal used to stand in for a quantity that is expensive to measure directly. |
Every term the collection defines is gathered in the glossary.