Clustered breakdowns
A Socratic walk-through of Clustered breakdowns — reasoned out one step at a time, not lectured.
The question we started with
THE QUESTION #Why do three machines fail in one week after months of quiet?
The workshop runs quietly for two months. Then on Tuesday a compressor trips, on Wednesday a pump seal goes, and on Friday a drive fails. Somebody says what everyone is thinking: "it never rains but it pours" — and then, usually, "something has changed."
That second sentence is the interesting one, because it is a causal claim smuggled in behind a feeling. Before we go hunting for the change, it is worth asking a duller question: what would independent, unrelated failures look like if we plotted them on a calendar? If the answer turns out to be "clumpy", then the cluster is not evidence of anything and the hunt starts from a false premise.
Reasoning it through
REASONING #Picture the failures as unrelated events, each with a small constant chance of happening on any given day — the standard model for wear-independent, random faults. Ask yourself what pattern you expect. Most people picture something roughly even: a failure, a gap, a failure, a similar gap.
Now check that intuition against the mechanics. If the events are independent, then after one failure the machine has no memory of it — tomorrow's chance is exactly what it was yesterday. So what is the most likely gap until the next failure? Not the average gap. The most likely gap is a short one, because the waiting time between independent events follows an exponential distribution, and that distribution peaks at zero. Short gaps are the single commonest outcome. Clumping is not a departure from randomness; it is randomness's own signature.
Put numbers on it, so the feeling has something to argue with. Say the shop averages 0.6 failures a week — quiet by most standards. Treating arrivals as Poisson, the chance that any given week carries three or more is about 2.3 per cent. Small. But there are 52 weeks, and the chance of seeing at least one such week somewhere in the year is roughly 70 per cent. So a three-in-a-week cluster is not a rare event to be explained; in a quiet year it is the expected event, and its absence would be more surprising.
Why does the intuition fail so reliably? Because we recall the cluster and never notice the six-week stretch that contained nothing — and because a genuinely even spacing would require the events to coordinate, which independent events by definition cannot do. The evenness we expect is the fingerprint of something scheduled, not something random.
Does this mean never look for a cause? No, and this is where the reasoning has to stay honest. The argument shows the cluster is not by itself evidence; it does not show there is no common cause. Real clusters often do have one: three machines fed from the same supply that took a voltage dip, three bearings from one bad batch, a heatwave, a coolant change, or a fitter who did the same thing wrong on all three. Those are testable, and the test is specific — do the three share a batch, a circuit, a technician, an environment? Ask that. What the Poisson calculation buys is a null hypothesis: it tells you how big a clump has to be before the clumping itself is worth treating as a signal, so you stop launching investigations into ordinary noise.
And there is a consequence that survives either way. A maintenance team sized for the average will be overwhelmed by the clusters, because the workload arrives lumpy even when its mean is comfortable. Capacity has to be sized against the queue, not the mean — which is why "we only average one failure a week" is a bad reason to keep one fitter.
The analogy
THE ANALOGY #Raindrops on a paving slab. The average is perfectly steady, but the slab does not wet evenly — it gets three spots almost touching, then a bare patch wide enough to look deliberate. Nobody suspects the sky of a conspiracy in that corner. The drops are simply not talking to each other, and unnegotiated arrivals always clump.
raindrops truly are independent, whereas machines in one shop often are not — they share a power supply, an operator, a spares batch and a season — so the clumping model is the baseline you compare against, not a proof that the three failures were unconnected.
Clarifying the model
THE MODEL #One refinement matters. The argument above assumes a constant failure rate, which suits the flat middle of an asset's life. It does not suit a fleet that is old, where wear-out makes the rate genuinely rise with age, nor a fleet just after commissioning, where infant mortality does the same at the other end. A cluster in an ageing fleet may well be the wear-out phase arriving on schedule — several units bought together reaching the end of the same design life within weeks of each other. That is a real common cause, and it is the reason installation dates are worth checking before the calculation is trusted.
The other correction is subtler: clustering feels like it demands explanation partly because we only ask the question after seeing the cluster. Choosing the week because it was bad and then asking how unlikely that week was will always overstate the case. The honest comparison is over the whole year, which is what the 70 per cent figure was for.
A picture of it
THE PICTURE #How to readthe bars are one illustrative run of five independent failures; the flat line is their own average of 0.63 a week — the same underlying rate that produced both the empty stretch and the week that looked like a crisis.
What became clearer
WHAT CLEARED #Independent failures arrive in clumps because short gaps between them are the most probable outcome, not the least. The cluster is therefore the wrong trigger for an investigation; the right triggers are a shared batch, circuit, technician or season — and a maintenance capacity sized for the lumps rather than the mean.
Where to go next
ONWARD #- How reliability engineers separate wear-out from constant-rate failure using Weibull plots of time to failure.
- Why queueing behaviour means a team at 85 per cent utilisation feels far busier than one at 70.
- Common-cause failure analysis, and how redundant equipment quietly shares a single weakness.
Key terms
TERMS #| Term | What it means |
|---|---|
| Poisson process | a model of independent events occurring at a constant average rate, in which the count in a fixed window follows the Poisson distribution. |
| Exponential distribution | the distribution of waiting times between such events; its most probable value is zero, which is why gaps cluster. |
| Common-cause failure | several failures traceable to one shared influence — a batch, a supply, a procedure, an environment. |
| Infant mortality and wear-out | the early and late life phases in which failure rate is genuinely not constant. |
Every term the collection defines is gathered in the glossary.