The most expensive maintenance mistake is applying one strategy uniformly. Time-based overhaul on an asset with random failure characteristics wastes remaining life and introduces infant mortality. Condition monitoring on a cheap, redundant, easily replaced component wastes analyst time. The correct answer is per-asset, and the sorting rule is not complicated.
The four strategies
- Run to failure — deliberate, not negligent. Correct when the failure consequence is low, the asset is redundant or cheap, and repair is fast.
- Preventive (time or usage based) — replace or service at a fixed interval. Correct when there is a genuine wear-out pattern with an identifiable useful life.
- Predictive (condition based) — monitor a parameter and act on the trend. Correct when a measurable degradation signal exists and the P-F interval is long enough to plan.
- Detective — periodic function testing of protective devices that are otherwise dormant. Correct for safety valves, trips, standby pumps and alarms.
The P-F interval decides whether prediction is even possible
P-F is the time between the point where a failure becomes detectable (P) and the point of functional failure (F). Your inspection frequency must be at most half the P-F interval, or you will miss failures.
| Technique | Typical P-F interval | Implied inspection frequency |
|---|---|---|
| Vibration analysis (bearings) | 1-9 months | Monthly |
| Oil analysis (gearboxes) | 1-6 months | Monthly to quarterly |
| Thermography (electrical) | 1-6 months | Quarterly |
| Ultrasound (leaks, early bearing) | 2-12 months | Monthly |
| Audible noise and human senses | Days to weeks | Every shift |
If a technique's P-F interval is shorter than your realistic response time, it produces alarms you cannot act on. That is worse than not monitoring, because it trains people to ignore alarms.
The age-related failure assumption is usually wrong
Classical reliability studies across several industries found that only about 11 per cent of failure modes are age-related; the remainder show random or infant-mortality patterns. The practical consequence is blunt: scheduled overhaul of a complex assembly frequently makes reliability worse, because reassembly introduces new defects. Reserve time-based replacement for components with a demonstrated wear-out mode — belts, filters, seals, brake linings, lubricant.
Criticality first
Score each asset on three axes and multiply: consequence of failure (safety, quality, output), probability of failure (history), and detectability. Rank the result. In a typical plant, 15 to 20 per cent of assets account for 80 per cent of downtime cost — those get condition monitoring and an RCM analysis. The bottom half gets run-to-failure with a spares policy, deliberately and in writing.
Measure the program, not the activity
| Metric | Meaning | Healthy direction |
|---|---|---|
| Planned maintenance percentage | Planned hours / total maintenance hours | Above 75 per cent |
| Schedule compliance | PM tasks done on time / scheduled | Above 90 per cent |
| MTBF | Operating time / number of failures | Rising |
| MTTR | Repair time / number of repairs | Falling |
| PM find rate | PMs that discover a real defect | 10-30 per cent |
PM find rate is the underused one. If your preventive tasks almost never find anything, the interval is too short and you are spending money to confirm that machines are fine. If they find something more than about a third of the time, the interval is too long and you are catching failures late.
Getting started without a budget
Two things cost nothing and change outcomes: recording failure causes in a consistent taxonomy so a Pareto is possible, and operator basic care — cleaning, inspection, lubrication, tightening. Autonomous maintenance routinely delivers more first-year improvement than an instrument programme, because it puts eyes on equipment daily.


Discussion0
Sign in to join the discussion.