Tuning to the test cycle
A Socratic walk-through of tuning to the test cycle — reasoned out one step at a time, not lectured.
The question we started with
THE QUESTION #Why do engines that pass every certification test still pollute heavily on real roads?
A diesel car is certified against a legal limit for oxides of nitrogen. It passes. It is sold, driven for years, and independent measurement on the road finds it emitting several times that limit — routinely, not occasionally. No warning light comes on. Nothing is broken.
The immediate explanation is fraud, and in at least one famous case it genuinely was. But the gap was measured across manufacturers and across models, including ones where no cheating was ever alleged. That is the interesting version of the question: how does a system full of honest engineers, testing honestly against a published standard, produce cars that reliably fail in the world the standard was written to protect?
Reasoning it through
REASONING #Start with what a regulator needs from a test. It must be reproducible — the same car must give the same number in a lab in Turin and a lab in Cologne, or the whole certification regime collapses into argument. It must be affordable enough that every model can be run through it. And it must be identical for every manufacturer, or it is not fair.
Now ask what those three requirements force. Reproducibility means fixing the variables: a defined speed trace, a defined ambient temperature, defined gear-change points, a rolling road rather than a real one. Fairness means publishing the procedure so everyone faces the same test. Affordability means the test is short.
So the regulator has, of necessity, produced a fully specified, published, finite driving profile. And that is the whole problem in one sentence, because a fully specified published profile is not a sample of driving. It is a particular sequence of operating points that everyone knows in advance.
Now put yourself on the other side and ask what the manufacturer is actually paid for. Not for clean air — nobody measures that at the point of sale. They are paid for passing the test, and separately for the things customers notice: fuel economy, throttle response, engine durability, service intervals. An engine control unit has dozens of tunable parameters, and many of them trade emissions against exactly those things. Recirculate more exhaust and you cut nitrogen oxides but foul the intake and lose power. Inject later and you cut them again but lose fuel economy.
So what is the rational calibration? Set the emissions-favouring parameters aggressively across the range of temperatures, speeds and loads the test visits — and set them for durability and economy everywhere else. Notice that no step in that reasoning requires anyone to lie. Every one of those choices can be defended in the room where it was made, often with a genuine engineering justification: exhaust recirculation really does cause deposits, and protecting the engine below a certain ambient temperature really is a legitimate concern. It is the aggregate that is dishonest, and no individual is holding the aggregate.
This is the ordinary shape of an incentive failure. The measure was chosen as a proxy for something we care about; the measure became the thing rewarded; and effort flowed to the proxy rather than the target. What makes the automotive case so stark is that the proxy is not merely imperfect, it is published and repeatable, which converts a soft measurement error into a precise optimisation target.
The scale of it is documented. Analysts at the ICCT tracked the divergence between the European type-approval fuel consumption figure and what drivers actually recorded, and found it widened from under ten percent in the early 2000s to around forty percent by the mid-2010s. Nothing suggests engines got worse over that period. What grew was the accumulated skill at the specific test.
And the fix tells you the diagnosis was right. Europe did not tighten the limit — it changed what gets measured, adding on-road testing with portable analysers on routes that are deliberately not fully specified in advance, alongside a more demanding laboratory cycle. Making the test partly unpredictable removes the thing that made it optimisable. Honestly, this is mitigation rather than cure: any test still has boundaries, and a boundary is still something that can be calibrated around.
The analogy
THE ANALOGY #Think of a school that is judged solely on its pupils' scores in one published exam. Teaching drifts toward the exam's format, its recurring question types, its marking scheme. Scores rise. Whether the pupils understand more is a separate question that nobody is being paid to answer.
a school's teachers can see they are narrowing the curriculum and can feel bad about it, whereas an engine calibration is distributed across hundreds of parameter maps and dozens of specialists, so the narrowing is nobody's decision and is visible only in the finished vehicle — which is why the automotive version is harder to catch and harder to blame.
Clarifying the model
THE MODEL #Three refinements.
First, the misconception worth dismantling: this is not primarily a story about cheating devices. Defeat devices are the extreme tail of a distribution whose bulk is entirely legal calibration toward a known target. Prosecuting the tail leaves the mechanism intact.
Second, the test being unrepresentative and the test being optimisable are different faults, and only the second is fatal. An unrepresentative but unpredictable test would still give roughly honest rankings between manufacturers. A representative but fully published test would still be gamed. Predictability is the load-bearing defect.
Third, and generalising: this pattern belongs to regulation as a whole, not to cars. Any published, repeatable compliance procedure that is cheaper to satisfy than the outcome it stands for will be satisfied rather than the outcome. The design question is not "is this test accurate?" but "what is the cheapest way to pass it, and does that path also deliver what we wanted?"
A picture of it
THE PICTURE #How to readEach bar is how far real-world fuel consumption sat above the certified laboratory figure for European cars of that vintage, following the ICCT's tracking of the divergence; treat the values as the reported trend rather than precise annual measurements. Read it as a record of accumulating skill at one fixed test, not of engines deteriorating — the laboratory number stayed compliant throughout while the thing it was standing in for drifted away underneath it.
What became clearer
WHAT CLEARED #The engines are not failing the test and they are not failing on the road by accident. They are succeeding at precisely what they were rewarded for, which was passing a published, repeatable procedure — and the road was never the thing being measured. The defect lives in the incentive, not in the machinery, which is why it survived every generation of tightened limits until the measurement itself was made unpredictable.
Where to go next
ONWARD #- Goodhart's law, and why proxy measures degrade the moment they become targets.
- How on-road testing with portable analysers changes what can be optimised.
- Whether outcome-based regulation, measuring fleet air quality rather than individual vehicles, dodges the problem or merely relocates it.
Key terms
TERMS #| Term | What it means |
|---|---|
| Test cycle | the fixed speed-versus-time profile a vehicle is driven through during certification. |
| Type approval | the regulatory sign-off that a vehicle model meets the applicable standards and may be sold. |
| Defeat device | a control strategy that detects test conditions and alters emissions behaviour accordingly; illegal, and narrower than the problem described here. |
| Real driving emissions testing | on-road measurement using portable analysers, introduced in Europe to break the predictability of the laboratory cycle. |
| Exhaust gas recirculation | routing some exhaust back into the intake to lower combustion temperature and hence nitrogen oxide formation, at a cost in deposits and power. |
Every term the collection defines is gathered in the glossary.