Search for "RCM vs predictive maintenance" and you will find page after page presenting the two as competing strategies, complete with comparison tables scoring one against the other on cost, sophistication and return. Those tables are a category error. You cannot compare them on cost or sophistication any more than you can compare a building code with a concrete pour. One is the reasoning process that decides what should be done to an asset; the other is one of the specific things that reasoning process might tell you to do. Once you see that, most of the confusion in this space dissolves, and something far more useful takes its place: a clear set of tests for whether a condition-based or predictive task is actually the right answer for a given failure mode.
The message up front: RCM is a decision framework; predictive maintenance is a task type. RCM does not compete with predictive maintenance, it decides whether a predictive task is valid for each failure mode, and it says no more often than vendors would like. The four tests it applies are detectability, a usable warning interval, technique capability, and a worthwhile action at the end. Fail any one of them and the honest answer is a different task entirely.
1. What RCM is, in one paragraph
Reliability-centred maintenance is a structured method for deciding what maintenance a physical asset needs in its current operating context. It starts from functions rather than equipment: what is this asset required to do, and to what standard? It then works out how each function can fail, what failure modes cause each functional failure, what happens when each one occurs, and only then what task, if any, is worth doing about it. The output is a maintenance policy per failure mode, not per asset. The criteria that define whether a process may legitimately be called RCM are published as SAE JA1011_202411 , "Evaluation Criteria for Reliability-Centered Maintenance (RCM) Processes"; a process that fails any of its criteria is not RCM regardless of what it is marketed as. SAE JA1012_201108 is the companion guide, now one generation behind the criteria document, and IEC 60300-3-11:2009 is the international application guide. For the full method, including the seven questions and the consequence categories, see the RCM introduction.
2. What predictive maintenance is, in one paragraph
Predictive maintenance is a task type. You measure some physical parameter that changes as a failure develops, you track it over time, and you intervene when the evidence says the failure is progressing toward functional failure. European maintenance terminology is genuinely helpful for pinning this down: EN 13306:2017, the CEN maintenance terminology standard, places predictive maintenance as a form of condition-based maintenance, which in turn sits under preventive maintenance alongside predetermined (time or usage based) tasks. That taxonomy matters because it tells you predictive maintenance is not a separate philosophy standing outside the preventive family, it is a particular way of triggering a preventive intervention. The supporting technical standards sit in the ISO condition monitoring family: ISO 17359:2018 as the programme-level general guidelines, ISO 13379-1:2025 for diagnostics (what is wrong and why) and ISO 13381-1:2025 for prognostics (how long is left). Both 2025 editions changed the series wording from "machines" to "machine systems", which is a small but telling shift toward treating the monitored thing as a system rather than a lump of rotating metal. For the techniques themselves, see the condition monitoring techniques guide, and for the delivery side, the predictive maintenance practitioner's guide.
3. The category difference, set out plainly
Here is the comparison table this article exists to replace. Note the first row, which is the one that actually resolves the question.
| Aspect | RCM | Predictive maintenance |
|---|---|---|
| What it is | A decision framework. A method for choosing maintenance policy. | A task type. One of the things a framework can choose. |
| Unit of analysis | The failure mode, reached through function and functional failure | The measurable condition parameter on a specific asset |
| Output | A task, an interval, and a justification, per failure mode. Sometimes the task is "none". | A monitoring regime, an alarm or trend threshold, and an intervention trigger |
| Governing documents | SAE JA1011_202411 (criteria), SAE JA1012_201108 (guide), IEC 60300-3-11:2009 | EN 13306:2017 (terminology), ISO 17359:2018, ISO 13379-1:2025, ISO 13381-1:2025 |
| Question answered | What is worth doing about this failure mode, and why? | Is this specific degradation progressing, and how fast? |
| Can it replace the other? | No. RCM without any condition-based options in its task menu is impoverished. | No. Predictive monitoring without an analysis behind it is instrumentation, not a strategy. |
| Typical failure when done badly | Analysis paralysis: months of workshops, no implemented change | Alert fatigue: dashboards full of signals nobody acts on |
The relationship runs one way. An RCM analysis reaches a point, for each failure mode, where it asks whether a condition-based task is applicable and worth doing. If the answer is yes, a predictive or condition-based task is the selected policy. If the answer is no, RCM moves on to other options. Predictive maintenance is downstream of the decision, never a rival to the process that makes it.
The test that settles it
Ask which one could produce the other as an output. An RCM analysis can produce a predictive monitoring programme. No amount of predictive monitoring produces an RCM analysis. That asymmetry is what makes them different categories rather than competing options.
The companion question, how RCM relates to traditional time-based preventive maintenance, is a genuinely different discussion and it has its own home: see RCM vs preventive maintenance. This article stays on the condition-based side of the task menu. If you want the wider strategy landscape rather than either comparison, preventive vs predictive vs reactive maintenance covers it.
4. The four conditions RCM requires before a predictive task is justified
This is the substance. RCM does not accept a condition-based task because it sounds modern. It applies tests, and a proposed monitoring regime has to pass all of them. Marketing material for predictive platforms routinely skips these tests, which is precisely why so many condition-monitoring programmes end up producing alerts nobody acts on. The four conditions, in the order I would apply them:
| Condition | The question to ask | What it means if the answer is no |
|---|---|---|
| a. A detectable degrading condition exists | Does this failure mode produce a measurable change before the function is lost, or does it arrive without warning? | There is nothing to monitor. No sensor helps. RCM must select a different task type or accept the failure. |
| b. The warning interval supports a practical inspection frequency | Is the interval between first detectability and functional failure long enough that a realistic monitoring frequency will catch it reliably? | The task is theoretically valid but practically useless. You will detect the fault after it has already caused the failure. |
| c. The chosen technique can actually see this failure mode | Does this specific measurement respond to this specific degradation mechanism, in this installation, at this stage? | You have monitoring that is blind to the thing you are worried about, and false confidence is worse than none. |
| d. A worthwhile action exists once detected | When the alert fires, is there something you can and will do that is meaningfully better than what you would have done anyway? | The monitoring generates information with no decision attached. Drop it, or fix the response capability first. |
5. Condition (a): is there anything to detect?
Every condition-based task rests on the assumption that failure is a process rather than an event. Something degrades, and the degradation is observable before the function is lost. That assumption holds for a great many failure modes: bearing surfaces spall and the vibration signature changes, insulation degrades and leakage current rises, a heat exchanger fouls and the approach temperature drifts, a gearbox wears and the oil carries the metal away with it. For those, a detectable condition genuinely exists and the first test is passed.
It does not hold universally, and this is where honest analysis differs from optimistic procurement. Failure modes driven by a random external event, a voltage transient that kills a control board, a foreign object drawn into an impeller, a fastener that fails in overload because someone over-torqued it during the last intervention, produce no antecedent degradation to observe. There was nothing wrong on Tuesday and the asset is down on Wednesday. Monitoring adds cost and generates no warning, because there was no warning available to capture.
The mistake I see most often here is analysing at the wrong level. Teams ask "is this pump's failure detectable?" when the answerable question is "is this pump's bearing-wear failure mode detectable?" A pump has many failure modes, some with rich warning signals and some with none at all. Monitoring the detectable ones is entirely rational, but it does not make the pump predictable, and a programme that sells itself internally as having made the pump safe will be embarrassed by the first failure mode it never covered. Working at failure-mode level is a discipline worth building; failure modes and how to analyse them sets out how.
6. Condition (b): is the warning interval long enough to use?
Detectability alone is not enough. The degradation has to develop slowly enough, relative to how often you can realistically look, that you catch it in time to act. This is the P-F interval logic: the span between the point at which a developing failure first becomes detectable and the point of functional failure. The monitoring interval has to sit comfortably inside that span, because a single observation at the wrong moment tells you nothing, and an inspection that arrives just after functional failure is not condition monitoring, it is a very expensive way of confirming a breakdown. The mechanics of this, including why the inspection interval has to be a fraction of the warning interval rather than equal to it, belong to the P-F curve explained, and the validity conditions in this article depend entirely on that model.
Two practical consequences follow. First, a short warning interval does not automatically rule out a condition-based task, it rules out a periodic inspection route and pushes you toward continuous monitoring, which is a different cost structure and a different integration problem. Continuous sensing has made some failure modes viable that were not viable under a monthly handheld route. Second, the warning interval is a property of the failure mode and the operating context, not a number you can look up. Duty cycle, load, ambient conditions and how hard the asset is run all shift it. I would not accept a vendor's generic figure for it, and I would not publish one here either, because the number that matters is the one your own failure history and your own trend data support.
Where this test is quietly failed
Routes get set at whatever frequency the team can staff, not at whatever frequency the failure mode requires. The route then runs for years, compliance looks excellent, and the programme still misses faults because the interval was never derived from the warning interval in the first place. If you inherit a condition-monitoring route, ask where its frequencies came from before you defend them.
7. Condition (c): can the technique see this failure mode?
This is the test most often skipped, because it requires technical judgement rather than a purchase order. A monitoring technique is not a general-purpose fault detector. Each one responds to particular physical changes, and a technique that is excellent for one failure mode can be effectively blind to another on the same machine.
Vibration analysis is superb for rotating-element faults, imbalance, misalignment, looseness and gear-mesh defects, because those produce energy at characteristic frequencies. It is far weaker on a slow internal erosion that does not change the machine's dynamic behaviour until late. Infrared thermography finds developing electrical connection faults and thermal anomalies extremely well, and tells you very little about a mechanical fault that has not yet started generating heat, or about anything behind a reflective or enclosed surface it cannot see. Oil analysis reads wear-surface degradation and contamination with long lead times, and says nothing at all about the electrical side. Ultrasound finds leaks, early lubrication distress and partial discharge, and is not a substitute for spectral vibration work.
The standards framework helps here more than people expect. ISO 17359:2018 sets out the programme-level logic of selecting measurement parameters against the failure modes you are trying to detect, rather than against the instruments you happen to own. ISO 13379-1:2025 covers diagnostics and ISO 13381-1:2025 covers prognostics, and keeping those two separate in your own head is useful: a technique may be perfectly capable of telling you something is wrong while being incapable of telling you how long you have. Where vibration severity assessment is involved, be careful with the reference standards. The ISO 10816 series has been superseded part by part by ISO 20816, and the migration is unfinished. ISO 10816-6:1995 for large reciprocating machines and ISO 10816-7:2009 for rotodynamic pumps remain the current documents for their scopes. Anyone telling you ISO 10816 has been wholesale replaced is working from a summary rather than the catalogue; check the ISO catalogue before you write a standard number into a specification.
The installation matters as much as the technique. A sensor mounted where the signal path to the fault is poor, a thermographic survey conducted through a closed panel door, an oil sample drawn from a dead leg rather than a live line: each is a technically correct technique rendered incapable by how it was deployed. Condition (c) is about the whole measurement chain, not the brochure specification of the instrument.
8. Condition (d): is there a worthwhile action at the end?
A condition-based task only earns its place if detecting the developing failure lets you do something better than you would have done without it. That sounds obvious and it is routinely ignored.
The clearest failure of this test is when the action on detection is identical to the action you would have taken on a schedule anyway, on an asset where the scheduled action is cheap and quick. If the response to a rising trend is "replace the filter element", and replacing the filter element is a twenty-minute job on a routine round, the monitoring has bought you nothing but a sensor to maintain. Do the scheduled task and spend the monitoring budget where the intervention is expensive or the downtime is consequential.
The subtler failure is organisational rather than technical. The monitoring works, the alert is correct, and nothing happens, because there is no spare, no window, no budget line, or no authority to take the asset off line on the strength of a trend chart. There are programmes where the engineering was sound and the entire value leaked away at this point. If an alert cannot become a scheduled, resourced work order in the system the technicians actually use, the loop does not close, and within a couple of months people stop looking at the dashboard. Condition (d) is as much a test of the maintenance organisation as of the failure mode.
The honest limitation of this whole approach
Applying four tests rigorously across a full asset register is a serious analytical effort, and classical RCM is expensive in exactly this way. Most organisations cannot afford it everywhere, and should not try. Run it properly on the failure modes where the consequence justifies the analysis, and use simpler judgement elsewhere. A framework applied everywhere shallowly is worse than one applied selectively with rigour.
9. The failure modes where predictive monitoring is simply the wrong answer
Pulling the four tests together gives a clear list of situations where RCM will decline a condition-based task. Naming them plainly is more useful than any positive case study.
- Failures with no detectable precursor. Random, event-driven or overload failures that arrive without antecedent degradation. Nothing to trend, so nothing to monitor.
- Warning intervals too short to work with. Degradation that runs from first detectability to functional failure faster than any practical monitoring regime can catch, and where continuous monitoring is not justified by the consequence.
- Failure modes the available technique cannot see. Either the physics does not produce a signal the instrument responds to, or the installation blocks the measurement. Monitoring here creates false assurance, which is actively harmful.
- Cases where the detected action equals the scheduled action. If you would do the same cheap task on a schedule anyway, the monitoring adds cost without adding a decision.
- Hidden failures with no on-condition signal. Protective devices and standby equipment whose failure is not evident in normal operation often cannot be trended at all, because the failed state looks identical to the working state until the function is called on.
What RCM selects instead, depending on which test failed and what the consequence is:
- A failure-finding task for hidden failures. You cannot monitor a standby pump's inability to start, so you periodically try to start it. This is not condition monitoring and it is not a time-based overhaul; it is a distinct task type whose whole purpose is to discover whether a hidden function has already failed. It is the single most under-implemented task type I encounter.
- A predetermined time or usage-based task where age-related wear genuinely dominates and a scheduled restoration or discard is defensible. The companion article covers when this is and is not sound.
- Redesign or a change of operating context where no task is adequate and the consequence is unacceptable. Adding redundancy, changing a material, relocating an asset out of a hostile environment, or removing the function altogether. RCM treats redesign as a legitimate output, not an admission of defeat.
- An explicit run-to-failure decision where the consequence is tolerable and no task is worth its cost. The word that matters is explicit. A documented, reasoned run-to-failure policy is a professional output. The same outcome arrived at by neglect is not, and the difference is entirely in whether anyone decided it.
Which of these applies depends heavily on consequence, which is why criticality work sits upstream of all of it. See equipment criticality analysis for how to establish the consequence picture before you start selecting tasks.
10. The honest relationship with technology: RCM answers have a shelf life
There is a genuine two-way tension here, and both sides of it deserve fair treatment.
On one side, technology changes what is detectable, and therefore changes RCM's answers. A failure mode that had no practical on-condition task ten years ago may have one now, because a sensor that was laboratory-grade is now cheap enough to leave permanently mounted, or because analytics can extract a signal from data that was previously just noise. That means an RCM analysis is not a permanent artefact. The functions and failure modes are reasonably stable; the answer to "is there an applicable and effective on-condition task" is not. An analysis performed years ago against the instrumentation of its time may be over-conservative today, recommending intrusive scheduled work for failure modes that could now be monitored instead. If your RCM documentation has never been revisited since it was written, its task selections are dated even if its failure analysis is not. Reviewing task selection when the detection option set materially changes is part of keeping the framework honest, and it is a much smaller job than the original analysis.
On the other side, and this is the direction the money usually flows, installing sensors without the analysis produces data rather than reliability. The instrumentation tells you what is changing. It cannot tell you which changes matter, what function is at risk, what the consequence of losing it is, or whether an action is available and worth taking. Those are the questions the framework exists to answer, and they do not answer themselves from a time-series database. A programme that buys the platform first and looks for the failure modes afterwards ends up with coverage decided by which assets were easy to instrument, which is not the same thing as which assets needed monitoring.
Both criticisms are fair
RCM conducted without current condition-monitoring knowledge is over-conservative and will over-prescribe scheduled intrusive work. Predictive technology deployed without analysis is unfocused and will over-instrument the accessible rather than the important. The competent position is neither camp: use the framework to decide, and keep its view of what is detectable up to date.
11. How they work together in practice
The productive arrangement is a loop, and it runs in a specific order.
- Establish consequence first. Criticality work identifies where analytical effort is justified. This is triage, not analysis.
- Analyse functions and failure modes on the assets that survive triage. Do this at failure-mode level or the later tests have nothing to bite on.
- Apply the four conditions to each candidate condition-based task. Write down the reasoning, including the rejections. The rejections are the most valuable part of the record, because they are what stops the same unjustified sensor proposal returning every budget cycle.
- Select the technique against the failure mode, not against the instrument inventory. Then check the measurement chain and the mounting, because the technique on paper and the technique as installed are not the same thing.
- Close the loop into the work management system. An alert has to become a real, scheduled, resourced work order, executed and closed out with a recorded finding. Any competent maintenance management system can carry this; the requirement is generic and the integration is the part that is usually underestimated.
- Feed findings back. Every intervention is evidence about whether the detection worked, how much warning it actually gave, and whether the threshold was set sensibly. This is how monitoring regimes get better, and it only happens if the closed work order records what was found rather than "completed".
- Revisit task selection when detection capability changes. Not the whole analysis, just the on-condition question. Make it a periodic review rather than a project.
For the broader engineering discipline this sits inside, including how reliability work relates to design, operations and data, see what reliability engineering actually is.
12. Deciding what you actually need
Because the two are not alternatives, "which should we do" is the wrong question. The useful questions are about sequencing and proportion.
- If you have monitoring but no analysis, and the symptom is alerts nobody acts on, your gap is the framework. Do not buy more sensors. Take your noisiest monitored assets and run the four conditions retrospectively against what you are already monitoring. Expect to switch some of it off, and expect that to be the highest-return work of the year.
- If you have analysis but it predates your instrumentation, revisit the on-condition question only. You are looking for scheduled intrusive tasks that could now become monitored tasks, and for failure modes previously marked undetectable that are no longer.
- If you have neither, start with consequence, not technology. Criticality, then failure modes on the critical few, then task selection. The first genuinely useful predictive deployment should come out of that sequence rather than ahead of it.
- If a vendor is comparing their platform to RCM, treat it as a signal about the vendor. A product can implement condition-based tasks. It cannot be an alternative to deciding which tasks are worth doing, and a pitch that claims otherwise is asking you to skip the reasoning.
- If someone insists you must choose, the category error has made it into the room. Restate it: the framework selects, the task executes. Then get back to the failure modes.
The idea to walk away with
RCM and predictive maintenance are not rivals and cannot be scored against each other, because one is the process that decides and the other is one of the things it may decide on. The value of understanding that is not terminological neatness, it is the four tests. A condition-based task is justified when a degrading condition is genuinely detectable, when the warning interval supports a workable monitoring frequency, when the chosen technique can actually see that specific failure mode as installed, and when a worthwhile action exists at the end. Fail any one, and the right answer is a failure-finding task, a scheduled task, a redesign, or an explicit and documented decision to let it run to failure.
Most disappointing condition-monitoring programmes did not fail on technology. They failed because nobody applied those four tests before the sensors were specified, and the programme ended up monitoring what was easy to monitor rather than what needed monitoring.
Final thoughts
The most useful thing about holding the categories apart is that it lowers the temperature. You stop arguing about whether RCM is old-fashioned or whether predictive maintenance is overhyped, and you start asking specific answerable questions about specific failure modes. Is there a signal? Is there time? Can this instrument see it? Will anyone act? Four questions, asked honestly, will do more for your reliability programme than any platform selection.
And if the answer to any of them is no, that is not a failure of the analysis. That is the analysis working. The whole point of a decision framework is that it is allowed to say no, and a framework that only ever approves the fashionable option is not deciding anything.
Disclosure
Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.
Reviewing a condition-monitoring programme?
Independent advisory on RCM task selection, whether your monitoring is pointed at the right failure modes, and closing the loop from alert to completed work order. 22+ years across utilities, oil and gas, manufacturing, government and facility operations. No sensor vendor margins, no reseller arrangements.
Book a conversationRelated reading: RCM vs preventive maintenance, RCM introduction, The P-F curve explained, Condition monitoring techniques, Predictive maintenance practitioner's guide, Preventive vs predictive vs reactive, Failure modes and how to analyse them, Equipment criticality analysis, What is reliability engineering.
Muhammad Abbas
CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.
Work with me