THIS EXPLANATION
THE ROOM
WRK·30 Work, Careers & Skilled Trades 6 MIN · 8 STATIONS

Preventive replacement

A Socratic walk-through of preventive replacement — reasoned out one step at a time, not lectured.

abcdefgh
a

The question we started with

THE QUESTION #

Why does a skilled worker replace a tool that is still working?

Watch a machinist pull a cutting insert that is still making good parts and something looks wasteful. The tool worked. It was thrown away anyway. The usual defence — "better safe than sorry" — is not an argument, since it would justify replacing everything every morning. So what actually decides it? And is there a case where replacing a working part is not merely cautious but provably a waste of money?

b

Reasoning it through

REASONING #

Begin with the thing that matters and is easy to skip past: not how long the part has lasted, but how likely it is to fail in the next hour, given that it has survived this long. Reliability engineers call that the hazard rate, and the whole question turns on whether it rises with age.

Suppose it does not. Suppose a component's chance of failing in the next hour is the same whether it is fresh out of the box or has run nine thousand hours. Then what have you bought by swapping it? The new one faces exactly the same chance of failing in the next hour as the old one did. You have paid for a part and a stoppage and purchased no reduction in risk at all. This is not a close call — for a constant hazard rate, preventive replacement is pure waste, and the arithmetic says so without any appeal to judgement. Many electronic components, and anything whose failures are triggered by external shocks rather than accumulated wear, behave close to this.

The next case is stranger. What if the hazard rate is falling — high at first, then settling down? That is the shape of manufacturing defects, bad installations, wrong torque, a seal nicked on assembly. An item that has run three months has proved it was not one of the bad ones. Replace it and you swap a survivor for an unproven unit at the riskiest point in its life. Here preventive replacement does not merely waste money; it raises the failure rate. This is why "we serviced it and then it broke" is a real pattern rather than bad luck.

Only in the third regime does replacement earn its keep: a hazard rate that increases with age, because something is accumulating — abrasion, fatigue cycles, corrosion, creep. A cutting edge dulling, a timing belt losing elasticity, a bearing spalling. Put the three regimes end to end and you get the shape called the bathtub curve: a falling stretch, a long flat stretch, and a rising tail.

So the first test is not "how old is it" but "does its failure rate rise with age at all?" — and for a great many parts the answer is no. The airline reliability studies of the 1960s and 70s that gave rise to reliability-centred maintenance found that only a minority of failure modes, roughly a tenth, showed a wear-out pattern a scheduled replacement could exploit.

The second test is the one workers apply intuitively. Even with a rising hazard rate, replacement pays only if failing unexpectedly costs more than stopping deliberately. A planned swap costs a part, a few minutes, and a slot you chose. An unplanned failure costs the part, plus whatever it damaged on its way out, plus scrapped work in progress, a crew standing idle, an expedited delivery — and the fact that it happened at the worst moment rather than one you picked. When that ratio is large, you replace early even at the price of throwing away good life. When it is near one, you run to failure on purpose, which is a strategy and not neglect.

c

The analogy

THE ANALOGY #
THE FIGURE

Think of a part as a die rolled once an hour, where a six means failure. A constant-hazard part is an ordinary die: it does not care that it has been rolled five hundred times, and a fresh die is no improvement whatever. A wearing part is a die that quietly gains an extra face marked six with every roll — and swapping that one out genuinely buys you back the odds you started with.

WHERE IT BREAKS DOWN

A die's odds are known exactly, while real hazard curves have to be inferred from thin failure data and often are not known at all; and a die gives no warning, whereas most wearing parts announce themselves through heat, vibration, swarf or noise — which is why condition monitoring, replacing on evidence rather than on the calendar, usually beats a fixed interval when the symptom is detectable.

d

Clarifying the model

THE MODEL #

First, "still working" is not evidence against replacement; in the wear-out regime every part is still working right up until it is not, and the entire point is to act before the event. But "still working" is a decisive argument in the other two regimes, and telling them apart is the skill.

Second, notice what the two tests do together. The hazard question decides whether replacement can reduce risk at all; the cost question decides whether the risk it removes is worth the certain cost of removing it. Fail the first and no cost ratio can rescue the schedule.

A caveat on the bathtub itself: it is a composite, useful for teaching and often wrong about a specific item. Real components show a whole zoo of curves, and many never leave the flat stretch. The disciplined version of this reasoning starts from the failure mode — what actually breaks and why — rather than from the shape of a textbook graph.

e

A picture of it

THE PICTURE #
Preventive replacement
Preventive replacement The curved line is the bathtub: read it left to right as one item ageing. The falling left stretch is early-life defects, where replacing makes things worse; the flat middle is the constant-hazard regime, where replacing changes nothing; the rising right tail is wear-out, the only stretch a scheduled replacement can exploit. The straight line is a different component whose hazard rate never rises at all -- and because it is flat everywhere, no replacement interval you choose for it will ever be better than running it to failure. {"generator":"[email protected]","source":"../Socrates/.diagram-cache/_src/preventive-replacement.md","sourceIndex":1,"sourceLine":4,"sourceHash":"c2d74fd3c7108025eb52306ddd1af1b79b9454698aa1d21e49dd780c0a72e0df","diagramType":"xychart","layoutVariant":"source","repairedDuplicateIds":[],"motion":"entrance-with-reduced-motion-fallback","presentation":"editorial","attempt":1,"viewBox":{"x":0,"y":0,"width":795,"height":668},"qa":{"passed":true,"findings":[]}} 0 1 2 3 4 5 6 7 8 9 Age in service 10 9 8 7 6 5 4 3 2 1 0 Chance of failing in the next hour

How to readThe curved line is the bathtub: read it left to right as one item ageing. The falling left stretch is early-life defects, where replacing makes things worse; the flat middle is the constant-hazard regime, where replacing changes nothing; the rising right tail is wear-out, the only stretch a scheduled replacement can exploit. The straight line is a different component whose hazard rate never rises at all — and because it is flat everywhere, no replacement interval you choose for it will ever be better than running it to failure.

f

What became clearer

WHAT CLEARED #
WHAT CLEARED

Preventive replacement is not caution, it is a bet that failure risk grows with age. Where it does, replacing early buys a genuine reduction in risk, worth the wasted life whenever an unplanned stoppage costs far more than a planned one. Where the hazard rate is flat, the bet is on nothing; where it is falling, the bet runs backwards. The worker who bins a working tool is trading a small certain cost for a large uncertain one — and only makes that trade on the rising stretch.

g

Where to go next

ONWARD #
  • How condition-based monitoring — vibration, oil analysis, thermography — replaces a calendar interval with evidence.
  • Why reliability-centred maintenance starts from failure consequences rather than from failure probability.
h

Key terms

TERMS #
TermWhat it means
Hazard ratethe probability that an item fails in the next interval given that it has survived to now; the quantity that decides whether age matters.
Bathtub curvethe composite shape of early-life, constant, and wear-out failure regimes plotted against age.
Condition-based maintenancereplacing on a measured symptom of impending failure rather than on elapsed time or cycles.

Every term the collection defines is gathered in the glossary.

Nearby on the shelf

4