Ask anyone who runs a building what good maintenance looks like and the answer is some version of "stay ahead of it." Inspect on a schedule, service on a schedule, replace before it breaks. Preventive is responsible; reactive is negligent. The whole industry is organised around the belief that more preventive maintenance is always better.
For most of the components in a building, that belief is backwards. The evidence, which comes from outside property entirely, from aviation and reliability engineering, is that the majority of equipment does not fail because it got old, and for anything that does not wear out with age, servicing it on a schedule does nothing useful and can actively make it less reliable. This article covers where that finding came from, why it means preventive maintenance is the wrong default for most assets, and how to tell the minority where it genuinely pays from the majority where it does not. One boundary first, stated plainly because it matters: safety-critical and code-mandated maintenance is not part of this argument, and nothing here suggests skipping it. More on that at the end.
The Finding That Broke the Assumption
The belief that equipment wears out on a predictable schedule, and should therefore be serviced on one, was the foundation of maintenance practice for most of the twentieth century. It was tested properly for the first time by the aviation industry, which had the strongest possible incentive to get it right.
In 1978, two engineers at United Airlines, Stanley Nowlan and Howard Heap, published a report titled Reliability-Centered Maintenance, sponsored by the US Department of Defense. It is the founding document of the discipline, and its central finding overturned the assumption the whole field had been built on. Studying failure data across aircraft components, they found that fixed-interval overhauls often did not improve reliability, and in many cases introduced failures that would not otherwise have happened. Most failure modes, it turned out, were not related to age at all.
The result is usually summarised in the six failure-pattern curves the report identified. The exact percentages vary between the original study and the ones that followed, so no single number should be treated as gospel, but the shape of the finding is consistent everywhere. Reliability engineers describe five of the six patterns as having a constant conditional probability of failure over their useful life, representing something like 83 to 97 percent of equipment depending on which study you take. A constant probability of failure means the thing is as likely to fail in any given month as in any other, regardless of how long it has been in service. In other words, it fails randomly, not with age.
Break the six patterns down and only three of them, accounting for roughly 15 percent of failures, show a genuine age-related wear-out pattern where failure becomes more likely as the component gets older. The other roughly 85 percent fail at a rate that does not rise with age, and a striking share of those are infant-mortality failures: things that fail early, from a defect, a bad installation, or a mistake made during service, then settle to a low constant rate if they survive the start.
Why This Makes Preventive Maintenance the Wrong Default
Scheduled preventive maintenance rests entirely on one premise: that failure becomes more likely as a component ages, so servicing or replacing it before it reaches that age heads off the failure. When that premise holds, preventive maintenance works. When it does not, the whole logic collapses, because for a component that fails randomly there is no age at which failure becomes more likely, and therefore no correct moment to preemptively service it. Doing so on a schedule cannot prevent a failure that was never going to be caused by age.
At best that work is wasted. At worst it is harmful, and this is the part that surprises people most. Every time a technician opens up a working piece of equipment to service it, they introduce the risk of infant mortality: a reassembly error, a disturbed connection, a part reinstalled slightly wrong, contamination let in. Because so many failures are early-life failures, intervening in something that was running fine can start the infant-mortality clock over again. The reliability engineers put it bluntly: time-based intrusive maintenance does not apply very often, and if arbitrarily applied it increases the risk of infant-mortality defects and failures, and increases maintenance costs. The scheduled service meant to prevent a failure can be the direct cause of the next one. For the majority of components, then, blanket preventive maintenance loses on both ends: it spends money and labour on work that cannot prevent the failure it targets, and the act of doing it introduces failures that would not otherwise have occurred.
The Two Questions That Sort Your Assets
None of this means maintenance is pointless. It means maintenance has to be matched to how a component actually fails, rather than applied uniformly on the assumption that everything wears out. Two questions do the sorting.
Does this component wear out with age? That is, does its probability of failure genuinely rise as it gets older, or is it roughly constant? Age-related items are things subject to steady physical wear: things that erode, corrode, fatigue, or abrade in service. Random-failure items are everything whose failure is triggered by an event rather than accumulated wear.
What happens when it fails? A failure that is dangerous, or that takes down something critical, or that is very expensive to fix reactively, is in a different class from one that is a minor, cheap, self-contained nuisance.
Put those together and you get four situations, each with a different correct strategy, and only one of the four is scheduled preventive maintenance.
| Fails with age (wear-out) | Fails randomly | |
|---|---|---|
| High consequence | Scheduled preventive maintenance or timed replacement. This is where PM genuinely earns its cost. | Condition monitoring. You cannot time it, so watch for the early signs and act on them. |
| Low consequence | Often still run to failure, if the reactive fix is cheap and safe. | Run to failure deliberately. Fix it when it breaks; scheduling anything is waste. |
The single quadrant where scheduled preventive maintenance is the right answer, high-consequence wear-out items, is real and important, but it is one of four, and in a typical building it covers a minority of the equipment. Everywhere else, the correct strategy is condition monitoring or deliberate run-to-failure, and "run to failure" here is not negligence. It is the financially and technically correct choice for a component that will not fail any sooner for being left alone and is cheap to fix when it does.
Why Property Operations Get This Wrong
If the reliability evidence is decades old and well established, why does property maintenance still run on blanket preventive schedules? Two reasons, both understandable.
The first is that a uniform schedule is easy to administer and a per-component strategy is not. "Service everything quarterly" fits on a calendar and can be handed to a crew. "Sort every asset class by failure pattern and consequence, then apply four different strategies" requires knowing how each thing actually fails, which is real analytical work. The uniform policy wins because it is simple, not because it is correct.
The second is cultural. Preventive maintenance feels virtuous and run-to-failure feels like giving up, so the instinct is always to add more preventive work, never less. A manager who cuts a PM task is exposed if anything later goes wrong, even if the task was worthless; a manager who keeps piling on preventive work is never blamed, even when the work is causing failures. The incentives push relentlessly toward over-maintenance, and the reliability evidence is the only thing that pushes back.
There is a specific cost to this in a property setting that connects to something operators already feel. Every hour a technician spends on unnecessary preventive work on a random-failure, low-consequence asset is an hour not available for the work that matters. Maintenance capacity is finite and, as anyone who has watched a queue back up knows, a fully loaded team develops long wait times fast. Over-maintaining the assets that do not need it consumes exactly the capacity that keeps response times short on everything else. Unnecessary preventive maintenance is not free even when it does no direct harm; it is spending your scarcest resource on the wrong assets.
What This Looks Like in a Building
Translating the framework into property terms, without pretending a blog can do the per-asset analysis for you:
The genuine wear-out items, where scheduled service earns its cost, tend to be things under continuous mechanical or thermal stress: HVAC components with moving parts and filters, pumps, belts, anything with a bearing under load, roofing and sealants exposed to constant weather. These have a real age curve, and catching them before the wear-out phase is worth doing.
The random-failure items, where scheduling accomplishes little, tend to be electrical and electronic components, controls, sensors, and anything that either works or does not with little warning from age. For these, condition monitoring where the failure has consequence, and deliberate run-to-failure where it does not, beats a service schedule.
And the infant-mortality reality argues for something operators rarely prioritise: the biggest single lever on failure is often not more maintenance but better installation and commissioning. Since so many failures are early-life failures caused by how something was put in, getting the installation right, by skilled people, to a high standard, prevents more failures than any amount of subsequent scheduled servicing. The reliability literature is emphatic on this point, and it is the opposite of where most maintenance budgets focus.
The honest move for any specific operation is to look at its own failure history: which assets actually fail, how often, and at what age. That data usually already exists in the completed work order record and is almost never analysed this way, which is the only reliable route from "we service everything on a schedule" to "we match the strategy to how each thing fails." The answer is in your own maintenance history rather than in a generic schedule, and a system that retains that history at the component level is what makes the analysis practical, RIOO among them.
The Boundary: Safety and Code
One firm line, because the argument above could be dangerously misread. None of this applies to maintenance that is legally required or safety-critical. Elevators, fire and life-safety systems, gas equipment, backflow prevention, and anything governed by a code-mandated inspection regime must be maintained on the required schedule regardless of its failure pattern, because the consequence of failure and the legal obligation override the economic optimisation entirely. Reliability-centered thinking is about where you have discretion. Where a code or a safety case removes that discretion, it is removed, and the schedule stands.
Conclusion
The instinct that more preventive maintenance is always better is one of the most durable ideas in operations, and one of the most thoroughly disproven. The aviation industry established decades ago, with the strongest imaginable motive to get it right, that most equipment does not fail from age, that servicing random-failure components on a schedule cannot prevent their failures, and that the act of intervention introduces failures of its own. The large majority of components do not benefit from the preventive work done on them, and some are made less reliable by it.
The goal is not more maintenance, and it is not less. It is maintenance matched to how each thing actually fails: scheduled service for the genuine wear-out items where failure has real consequence, condition monitoring for the high-consequence random failures you cannot time, and deliberate run-to-failure for the low-consequence remainder. Getting that match right costs less than a blanket schedule and produces more reliable equipment, which is the rarest kind of improvement: cheaper and better at once.
The reason it is rare in practice is not that the evidence is obscure. It is that "do more preventive maintenance" feels responsible and "let that one run until it breaks" feels negligent, when for most of the assets in a building the second is the correct engineering decision and the first is quietly wasting money and inducing the failures it was meant to prevent.
FAQs
1. Is preventive maintenance always better than reactive maintenance?
No. Preventive maintenance only reduces failures for components that genuinely wear out with age, so that servicing them before that age heads off the failure. The reliability research originating with Nowlan and Heap's 1978 report found that the majority of components do not fail from age; they fail at a roughly constant rate regardless of age. For those, scheduled preventive work cannot prevent the failure it targets and can introduce new failures through the disturbance of servicing.
2. What is reliability-centered maintenance?
Reliability-centered maintenance is a discipline, developed in aviation and formalised in a 1978 US Department of Defense report, that matches maintenance strategy to how a component actually fails rather than applying a uniform schedule. It distinguishes age-related wear-out failures, where timed service or replacement works, from random failures, where condition monitoring or deliberate run-to-failure is more effective, and weighs each against the consequence of failure.
3. Does run-to-failure mean neglecting maintenance?
No. Run-to-failure is a deliberate strategy for components that fail randomly rather than with age and whose failure is low-consequence and cheap to fix. For these, there is no age at which preemptive service would help, so scheduling work is wasted effort. Choosing to fix them when they fail, rather than servicing them on a calendar, is the technically and financially correct decision, not negligence.
4. Can preventive maintenance actually cause failures?
Yes, and this is one of the central findings of the reliability research. Opening up working equipment to service it introduces the risk of reassembly errors, disturbed connections, contamination, and other faults, which show up as early-life or infant-mortality failures. Because a large share of all failures are early-life failures, unnecessary intervention in equipment that was running correctly can cause the next failure rather than prevent it.
5. Does this apply to safety and code-required maintenance?
No. Maintenance that is legally mandated or safety-critical, such as for elevators, fire and life-safety systems, gas equipment, and backflow prevention, must be performed on the required schedule regardless of failure pattern. Reliability-centered thinking applies only where the operator has discretion over the maintenance strategy. Where a code or safety obligation removes that discretion, the required schedule governs.