Drift in accepted values of constants
A Socratic walk-through of Drift in accepted values of constants — reasoned out one step at a time, not lectured.
The question we started with
THE QUESTION #Why do successive measurements of a constant drift toward the previous answer instead of scattering around the truth?
Plot the published values of a physical constant against the year they were measured, and you expect a cloud: points scattered around the true value, tightening as instruments improve. What you often see instead is a staircase. Values sit close to one another for a decade, then march steadily in one direction, each new result landing just a little past the last.
Feynman pointed at exactly this in his 1974 Caltech commencement address, using Millikan's electron charge as the case. Millikan's value was slightly too low, because he had used a poor figure for the viscosity of air. The measurements that followed did not jump to the correct value. They crept toward it. Why would a sequence of independent experiments do that?
Reasoning it through
REASONING #Start by asking what "independent" is doing in that sentence. The apparatus is independent. The observers are independent. But is the decision to stop independent?
Consider what actually happens in a laboratory. You run the experiment, reduce the data, and get a number. It disagrees with the accepted value by three times your quoted uncertainty. What do you do next?
You do not publish. You go looking for what went wrong — and you are right to. A three-sigma discrepancy is far more often a systematic error in your own setup than a revolution in physics. You check the calibration, you re-derive a correction, you find that a temperature coefficient was applied with the wrong sign. You fix it. The number moves closer. Now you publish.
Now run the counterfactual, because this is the step that matters. Suppose your first number had landed right on the accepted value. Would you have gone hunting for that same sign error?
Almost certainly not. And there is the whole mechanism. The error hunt is real, competent, and honest — but it is triggered by disagreement and called off by agreement. The stopping rule is conditioned on the outcome. Every individual step is good practice; the sequence is biased.
Notice that no one has to be dishonest for this to work, which is why the effect is so durable. What is doing the work is not deception but incentive: the cost of publishing a discrepant number is high — you will be asked to explain it, your systematics will be scrutinised, you may be wrong in public — while the cost of publishing an agreeing number is nearly zero. Asymmetric costs applied to a symmetric search produce a systematically asymmetric result.
Two further pressures push the same way. Referees and collaborators apply the same asymmetry from outside; a paper reporting an anomaly gets harder questions than one reporting a confirmation. And there is the file-drawer effect: discrepant results that survive scrutiny may still be quietly shelved as "probably something we do not understand," so the published record is not the record of what was measured.
There is a diagnostic, and it is the reason we can be confident this is happening rather than merely suspect it. If a sequence of measurements were genuinely independent draws around a true value, successive results should differ from one another by roughly the size of their combined uncertainties. In these staircase sequences they do not. They agree with each other far too well for their quoted error bars, and then the whole cluster later turns out to be displaced from the modern value by many of those error bars. Too little scatter, followed by a large jump, is the signature. The Particle Data Group's own reviews discuss this history openly and show such trend plots for particle properties.
The analogy
THE ANALOGY #Think of a jury that deliberates by having each member state a verdict aloud in turn. Every juror is honest and reasons carefully. But the second juror, hearing the first, checks their own reasoning harder when they disagree than when they agree — and finds, honestly, a flaw. By the twelfth juror the room is unanimous, and the unanimity is nearly worthless as evidence, because eleven of the twelve opinions were partly determined by the first. Ask them all to write their verdict privately before anyone speaks and you get a very different distribution.
jurors are weighing the same fixed evidence, whereas each experiment genuinely produces new evidence — so the drift is not pure conformity, and better apparatus really does move the value toward the truth, just more slowly and more smoothly than it should.
Clarifying the model
THE MODEL #It is tempting to read all this as a charge of misconduct. It is not. The clearest version of the problem involves no one behaving badly at any point; it is a property of the procedure, and it survives replacing every person in it.
Nor does it mean published constants are untrustworthy. It means their quoted uncertainties are, in this specific way, optimistic — the scatter of the community's results underrepresents the real uncertainty, because the results are correlated through the literature rather than independent. This is precisely why some ongoing discrepancies are taken seriously as physics rather than dismissed as sloppiness: the discrepant measurements are ones where the groups deliberately insulated themselves from the expected answer.
Which points at the fix. Blind analysis: the analysis pipeline is written and frozen while the final number is hidden from the analysts — offset by an unknown constant, or with a random fraction of the data withheld — and unblinded only once the systematics are settled and everyone has agreed to publish whatever appears. Now the error hunt cannot be conditioned on the answer, because nobody knows the answer while hunting. Blinding is standard in particle physics and increasingly common elsewhere. Pre-registration of analysis plans in other fields is the same idea wearing different clothes.
The honest caveat: how large the residual drift is in any particular constant is a matter of judgement, not measurement. You cannot compute the size of a bias whose mechanism is "searches that were not performed."
A picture of it
THE PICTURE #How to readFollow the two exits from the top gate. The agreeing branch runs straight to publication with no scrutiny; the disagreeing branch loops through an error hunt that can only move the result toward the accepted value or remove it from the record entirely. The dotted back-edge closes the loop — today's published value becomes tomorrow's comparison, so the asymmetry compounds.
What became clearer
WHAT CLEARED #A biased result does not require a biased person. Here the bias lives entirely in a stopping rule — keep looking while you disagree, stop when you agree — and that single asymmetry is enough to turn a sequence of honest, competent measurements into a slow procession behind whoever measured first.
Where to go next
ONWARD #- Blind analysis in practice: how a collaboration adds a secret offset and what it gives up in exchange.
- Whether the same asymmetric stopping rule explains parts of the replication crisis outside physics.
- The proton radius and neutron lifetime puzzles, where two methods disagree persistently and neither camp is drifting toward the other.
Key terms
TERMS #| Term | What it means |
|---|---|
| Bandwagon effect in measurement | the tendency of successive published values to cluster near the previously accepted one rather than around the truth. |
| Systematic error | an offset that shifts every reading the same way and cannot be reduced by repeating the measurement. |
| Blind analysis | hiding the final result from the analysts until the analysis is frozen, so error hunting cannot depend on the answer. |
| File-drawer effect | the distortion of the literature caused by results that are measured but never published. |
| Particle Data Group | the international collaboration that compiles and reviews accepted values for particle properties. |
Every term the collection defines is gathered in the glossary.