mail@mabbaz.com Abu Dhabi, UAE

Predictive Maintenance · Reliability · CMMS / EAM

Predictive Maintenance: A Practitioner's Guide

Predictive maintenance is the most talked-about idea in maintenance and the least well understood. This is the category guide: what predictive maintenance actually means, how it differs from preventive and condition-based maintenance, where the return genuinely comes from, the prerequisites almost every organisation underestimates, and a realistic maturity path from where you are now to a program that works.

Muhammad Abbas September 24, 2026 ~22 min read

Ask ten maintenance professionals to define predictive maintenance and you will get ten answers, most of them describing something else. Some will describe a monthly vibration route, which is condition monitoring. Some will describe a temperature alarm, which is condition-based maintenance. Some will describe a machine learning platform they were sold and have never switched on. The confusion is not the reader's fault: the term has been stretched by two decades of marketing until it covers everything from a thermal camera to an artificial intelligence claim. This guide puts the definition back in its place, then walks through what it actually takes to do predictive maintenance in a real operation, in the order the work has to happen.

The message up front: predictive maintenance is not software and not sensors. It is a decision-making capability that turns evidence of a developing fault into a planned intervention, and it rests on four prerequisites most organisations skip: a criticality-ranked asset register, maintenance history clean enough to trust, an understanding of how each asset class actually fails, and somebody whose job it is to act on an alert. Buy the prerequisites first. The technology is the easy part and the cheap part.

1. What predictive maintenance is: the plain definition

Predictive maintenance, often abbreviated PdM, is a maintenance strategy in which the timing of intervention is decided by evidence of the asset's actual condition and its projected trajectory, rather than by a calendar, a meter reading, or the asset having already broken.

That definition has three working parts, and each one carries weight:

  • Evidence of actual condition. Something is measured on the real asset: vibration, temperature, oil chemistry, acoustic emission, motor current, or a process parameter such as discharge pressure. Not an assumption about how the asset is probably doing, an actual measurement.
  • A projected trajectory. The measurement is compared against history and against normal, so the question being answered is not only "is this asset unwell today" but "where is this heading, and how quickly." This is the part that makes it predictive rather than reactive to a threshold.
  • A decision about timing. The output is a maintenance decision: continue monitoring, plan an intervention in the next window, or act now. If the output is a chart that nobody converts into a work order, the strategy is not in operation, however good the technology is.

Practitioners sometimes shorten this to a sentence I find useful: predictive maintenance is maintenance scheduled by the asset rather than by the planner. The planner still plans, but the trigger comes from the equipment's own behaviour.

The one-line test

If you removed the monitoring tomorrow, would any maintenance decision change? If the answer is no, you are collecting condition data, not running predictive maintenance. That single question separates real programs from expensive instrumentation projects faster than any maturity assessment.

2. What predictive maintenance is not

Because the term is so loosely used, it is worth being explicit about the things that get labelled predictive maintenance and are not.

  • It is not condition monitoring on its own. Condition monitoring is the measurement activity. Predictive maintenance is what you do with the measurement. A site can run a thorough vibration route for years and remain entirely reactive if the readings are filed and not acted on.
  • It is not a threshold alarm. "Alert when bearing temperature exceeds 85 degrees" is condition-based maintenance, and a perfectly respectable thing to have. It tells you the asset is already degraded. Prediction is about seeing the degradation begin.
  • It is not artificial intelligence. Machine learning can strengthen predictive maintenance and in some settings makes it possible at all, but the strategy long predates it. A great deal of effective predictive maintenance is done with trend charts, engineering judgement and a spreadsheet. The algorithm is optional; the discipline is not.
  • It is not a replacement for preventive maintenance. Statutory inspections, lubrication, filter changes, calibration and regulatory testing do not go away because you installed sensors. Predictive maintenance replaces a specific category of preventive task: the intrusive, fixed-interval overhaul or replacement performed on assets that were probably still healthy. Everything else stays. The complete guide to preventive maintenance sets out what that remaining programme should look like.
  • It is not a plant-wide destination. This is the misconception that costs the most money. Predictive maintenance is a strategy you assign to selected assets, not a level the whole site graduates to. On most asset registers it earns its place on a minority of the equipment, and that is the correct outcome, not a half-finished project.

3. How it differs from preventive, reactive and condition-based maintenance

The four strategies are best understood by what triggers the work. That single distinction removes most of the confusion.

  • Reactive (run to failure): the trigger is the failure itself. Entirely appropriate for cheap, non-critical, quickly replaced items where the consequence of failure is inconvenience.
  • Preventive (time or meter-based): the trigger is elapsed time or accumulated usage. Predictable, auditable, easy to resource, and blind to whether the asset needed the attention.
  • Condition-based: the trigger is a measured condition crossing a defined limit. Better targeted than preventive, but by the time the limit is crossed the fault is usually well established.
  • Predictive: the trigger is a trend indicating a fault is developing, well before any limit is reached. Earliest warning, widest planning window, highest cost and highest skill requirement.

Condition-based and predictive maintenance sit on a continuum rather than either side of a hard line, and most working programs operate somewhere between them. That is not a failure of ambition. Condition-based maintenance delivers a large share of the achievable benefit for a fraction of the effort, and I would rather see an organisation running disciplined condition-based maintenance on fifty assets than a stalled predictive pilot on three. For a fuller side-by-side treatment of the strategies, see preventive vs predictive vs reactive maintenance, and for the decision framework that assigns a strategy per failure mode, reliability-centred maintenance.

The mechanics of how early detection actually works, the interval between a detectable defect and functional failure, how remaining useful life is estimated, and how the models behind it are built, are covered in depth in the companion guide on predictive maintenance and failure prediction. This article deliberately stays at the strategy layer and defers the prediction mechanics there.

4. The business case: where the return actually comes from

Most predictive maintenance business cases are built on the wrong benefit. They lead with reduced maintenance labour, which is the smallest and least reliable saving, and bury the benefit that actually carries the case.

The return, in the order I would rank it for a typical industrial or large facilities operation:

  • Avoided consequence of failure. Not the repair cost, the consequence: lost production, a service interruption, a regulatory breach, a safety event, a reputational hit. On a genuinely critical asset this dwarfs every other line in the case, and it is why criticality ranking has to come before the sensor selection.
  • Avoided secondary damage. A bearing replaced when the vibration signature first changes is a modest job. The same bearing run to destruction can take the shaft, the seal and the coupling with it, and turn a planned half-day into a multi-week rebuild with long-lead parts.
  • Converting emergency work into planned work. Planned work is cheaper per hour, safer, better prepared, does not consume overtime, and does not derail everything else in the schedule. Shifting the planned-to-reactive ratio is where a maintenance department's efficiency genuinely improves.
  • Eliminating unnecessary intrusive work. Opening a healthy machine carries its own risk. Infant mortality after intrusive maintenance is a real and well-documented phenomenon, and stopping the fixed-interval strip-down of equipment that was working is both a saving and a reliability gain.
  • Better parts and resource staging. Knowing three weeks out that a component will need replacing lets you order rather than expedite, and schedule a crew rather than scramble one.
  • Deferred capital replacement. Evidence of actual condition supports keeping an asset in service past a nominal design life, or replacing it earlier, on data rather than on age. Over a large register that is a material capital planning benefit.

Against that sits a cost base that vendors are quieter about: sensors and installation, the data platform, integration into the maintenance system of record, analyst or reliability engineering time, and the ongoing tuning that stops the whole thing decaying into noise. The detailed return modelling, including how to build a defensible case per asset, is handled in the failure prediction guide.

Where the business case falls apart

If your organisation cannot currently measure unplanned downtime, cannot separate planned from reactive work orders, and has no cost attached to an hour of asset outage, you cannot build a real predictive maintenance business case and you cannot prove the programme worked afterwards. Fix the measurement first. A programme with no baseline is a programme that will be cancelled in the next budget round regardless of whether it succeeded, because nobody will be able to demonstrate that it did.

5. The four prerequisites most organisations underestimate

This is the section I would keep if I had to delete every other one. The reason predictive maintenance programmes disappoint is almost never the sensor or the algorithm. It is that one of these four foundations was missing and nobody checked.

Prerequisite 1: a criticality-ranked asset register. You cannot decide where predictive maintenance belongs without knowing which assets matter and why. Criticality is not a gut ranking by the maintenance manager; it is a structured assessment of consequence across safety, environment, operations, cost and compliance. Without it, sensor budget gets allocated politically, to whichever asset caused the most recent embarrassment. With it, allocation follows consequence. The method is set out in asset criticality classification, and the register itself has to live in your CMMS or EAM rather than in a spreadsheet, or it will be out of date within a quarter.

Prerequisite 2: maintenance history you can trust. Predictive maintenance needs to know how your assets have actually failed. That information should already exist in years of closed work orders, and in most organisations it is unusable: failure codes applied inconsistently or not at all, downtime not recorded, causes left blank, completion notes that say "rectified". The uncomfortable arithmetic is that a site with five years of poor history is in a worse starting position than a site with one year of disciplined history. Cleaning this up is unglamorous, takes twelve to eighteen months of consistent closeout discipline, and is non-negotiable. The structural fix starts with getting work order types right so planned, reactive and condition-triggered work are distinguishable at all.

Prerequisite 3: an understanding of failure modes. You do not monitor an asset, you monitor a failure mode. A centrifugal pump can fail by bearing wear, seal failure, impeller erosion, cavitation damage, coupling misalignment or motor winding breakdown, and those require different measurements. If you do not know which failure modes dominate on your equipment and in your operating environment, the sensor selection is guesswork. This is what a failure mode and effects analysis is for, and a lightweight version on your top asset classes is usually enough to direct the monitoring strategy correctly.

Prerequisite 4: somebody whose job it is to act. The most consistently fatal gap. An alert with no owner is noise. There must be a named role, with time allocated, responsible for reviewing condition data, deciding what it means, and raising the work order. In smaller operations this is a portion of a planner's or supervisor's week. In larger ones it is a reliability engineer. What it cannot be is "everybody", or an extra duty added to a fully loaded supervisor. I have watched more programmes die from this than from every technical cause combined: the technology worked, the alerts fired, and nobody had the hours to look at them.

A readiness check before you spend anything

Can you produce a criticality-ranked list of your top fifty assets? Can you pull twelve months of failure history for one asset class and have it make sense? Can you name the dominant failure modes on that class? Can you name the person who will act on an alert next Tuesday? Four yeses means you are ready to pilot. Three or fewer means the next project is the missing prerequisite, not the sensors.

6. The condition monitoring techniques, at a glance

These are the measurement technologies that feed predictive maintenance. None of them is new; reliability engineers have used most of them for decades with handheld instruments. What changed is that continuous, permanently installed versions became affordable. The practitioner's task is matching technique to failure mode, not deploying everything everywhere.

Technique What it detects Best suited to Typical warning Relative cost
Vibration analysis Bearing wear, imbalance, misalignment, looseness, gear and blade defects Pumps, motors, fans, compressors, gearboxes, turbines Weeks to months Medium to high
Infrared thermography Loose or corroded electrical connections, overload, failing bearings, blocked cooling, insulation loss Switchgear, distribution boards, motors, steam systems, building envelope Days to months Low to medium
Oil and lubricant analysis Wear metals, viscosity change, water and particle contamination, additive depletion Gearboxes, engines, hydraulics, large bearing housings, transformers Weeks to months, often the longest lead time Low per sample
Airborne and structural ultrasound Compressed air and gas leaks, steam trap failure, early bearing lubrication faults, partial discharge, valve pass-through Compressed air networks, steam systems, switchgear, slow-speed bearings Days to weeks Low
Motor current signature analysis Broken rotor bars, air-gap eccentricity, stator faults, some driven-equipment faults Medium and large electric motors, especially inaccessible ones Weeks Medium
Motor insulation testing Winding insulation degradation, moisture ingress, contamination HV and critical LV motors, generators Months Low, but usually offline
Process parameter trending Efficiency loss, fouling, wear, blockage, control drift Almost anything already instrumented for control Weeks to months Very low, data already exists
Non-destructive testing Wall thinning, corrosion, cracking, weld defects Pressure vessels, pipework, structures, tanks Months to years High per inspection

The cheapest entry point on that table is the last-but-one row. Process parameter trending uses data your building management system, SCADA or control historian is already recording. A chiller whose approach temperature is creeping up at constant load, or a pump whose discharge pressure is falling at constant speed, is reporting a developing fault with no new hardware required. I would always exhaust that source before specifying a single new sensor, because it is free, it is already integrated, and it builds the analysis habit that the rest of the programme depends on.

For the formal framing of condition monitoring and diagnostics of machines, the ISO 17359 family is the reference standard, and the ISO 55001 asset management standard provides the wider management-system context that a monitoring programme should sit inside.

7. The data and sensor layer

Once you move past handheld routes into continuous monitoring, you are building a small data architecture whether you intended to or not. Understanding its layers keeps you from being sold a black box.

Sensing layer     (sensors, existing instrumentation, manual routes)
  ↓
Transport layer   (wireless gateway, fieldbus, OPC UA, MQTT)
  ↓
Storage layer     (time-series historian, data platform)
  ↓
Analysis layer    (thresholds, trends, anomaly models)
  ↓
Action layer      (CMMS / EAM work order, schedule, closeout)
  ↓
Feedback loop    (what was actually found, back into the analysis layer)

Two observations from implementation work. First, the action layer is chronically underspecified. Proposals describe sensors and dashboards in detail and then say "integrates with your CMMS" in a single line. That line is where the programme lives or dies: an alert has to become a work order, with the right type, the right priority, the right asset reference and the right trade, inside the system technicians already use. If a technician has to open a second application to find out what needs doing, adoption will not survive the first busy month. The integration patterns are covered in IoT integration with a CMMS.

Second, the feedback loop is almost always missing. When a predicted fault is investigated, what the technician actually found should be captured against the alert. Was the prediction right? Was it early, late, or wrong? Without that, thresholds never get tuned, models never improve, and the programme cannot answer the only question management will ask, which is how often the predictions are correct.

The false alarm problem

Set thresholds tight and you generate alerts on healthy equipment. Technicians investigate, find nothing, and within a few cycles they stop taking the alerts seriously. Once credibility is lost it is extremely hard to recover, because the next real warning arrives into an audience that has learned to ignore this source. Start deliberately conservative, accept that you will miss some early indications in the first year, and earn trust before you tighten sensitivity. A programme with a reputation for crying wolf is worse than no programme, because it has also consumed the budget.

8. Where machine learning genuinely helps

Machine learning has a real place in predictive maintenance, and it is narrower and more specific than the market suggests. At the category level, three honest statements cover it.

It helps most where the signal is multivariate. A human or a simple threshold handles one rising trend well. What people handle badly is a pattern across fifteen parameters where no single one is out of range but the combination is abnormal. That is exactly what anomaly detection is good at, and it is the application I would recommend first because it does not require a labelled history of past failures, only a reasonable picture of normal operation. See AI anomaly detection for early fault warning for how that works in practice.

It helps most where you have fleet scale. Two hundred similar pumps with comparable duty give a model something to learn from. Three unique bespoke machines do not, no matter how much data each one produces. Fleet homogeneity matters more than data volume.

It helps least as a substitute for the prerequisites. Machine learning consumes data quality, it does not create it. A model trained on inconsistent failure coding will produce confident, well-presented nonsense, which is more dangerous than no model at all.

The model families, how remaining useful life is actually estimated, and where each approach breaks down are covered in the failure prediction guide, with the pattern-recognition techniques in deep learning for maintenance pattern recognition. For the purposes of this pillar the strategic point is simply this: start with physics, thresholds and trends, which are transparent and which engineers will trust, and earn the right to complexity on the specific assets where simpler methods have demonstrably run out.

9. The organisational change nobody budgets for

Predictive maintenance changes how a maintenance department works, and the change is larger than the technology. This is consistently the least-planned part of any programme I am asked to review.

  • The planning horizon lengthens. Predictive alerts arrive weeks ahead of need, which only helps if you have a planning and scheduling process capable of absorbing work into a future window. Departments that operate week to week have nowhere to put a three-week-out prediction, so it either gets done immediately, wasting the advantage, or forgotten.
  • Work origination changes. Work orders start appearing from a source that is neither a PM schedule nor a fault report. That needs its own work order type, its own priority logic, and its own place in the weekly schedule review, or it will be treated as an interruption.
  • New skills are required. Somebody has to interpret a vibration spectrum or an oil report. You can buy that as a service, develop it internally, or get it from the sensor vendor, but you cannot skip it. Raw data with no interpretation capability is a subscription with no output.
  • Trust has to be built with technicians. A technician sent to a machine that sounds fine, on the strength of an algorithm, is being asked to accept an unfamiliar authority. Show them the finding when the strip-down confirms the prediction. Nothing converts scepticism faster than being proved right in front of the people doing the work.
  • The success measures change. A department judged on PM compliance and work order closure will resist a strategy that reduces PM count. Metrics have to move toward reliability outcomes: unplanned downtime, mean time between failures, the planned-to-reactive ratio, schedule compliance. The KPI framework guide sets out a workable set.

If I had to choose between a technically excellent programme in an organisation that has not made these changes, and a technically modest one in an organisation that has, I would back the second every time. The measurement technology is a commodity. The operating discipline is not.

10. A realistic maturity path

Nobody goes from reactive to fleet-wide predictive maintenance in a budget cycle, and the attempts to do so are where most of the wasted money sits. The path below is the sequence I would advise, with an honest view of how long each stage takes and what it actually requires. Stages are cumulative: you do not leave the earlier ones behind.

Stage What is in place What you do next Realistic duration Spend
0. Reactive Work is triggered by breakdowns. Asset register incomplete, no reliable history. Build the asset register, adopt a CMMS properly, enforce work order closeout with failure coding. 6 to 12 months Low, mostly effort
1. Planned PM schedules running, compliance measured, history accumulating, basic KPIs reported. Rank assets by criticality. Review PM content for value. Analyse the failure history you now have. 6 to 12 months Low
2. Targeted Criticality ranking complete, dominant failure modes understood on top asset classes. Start manual condition monitoring routes on critical assets. Thermography and oil sampling first, they are cheapest. 6 to 9 months Low to medium
3. Condition-based Routes running, readings trended, condition findings raising real work orders in the CMMS. Instrument the highest-consequence assets continuously. Integrate alerts into the CMMS. Name the owner of the alert queue. 9 to 18 months Medium
4. Predictive Continuous monitoring on critical assets, trend-based intervention decisions, feedback loop closing. Add analytics where the data supports it. Track prediction accuracy. Retire PM tasks the monitoring has made redundant. 12 to 24 months Medium to high
5. Reliability-led Strategy assigned per failure mode, predictive used selectively, reliability outcomes trending measurably. Feed findings into design, procurement and specification. Attack the causes, not just the symptoms. Continuous Sustaining

Three things to notice. Stages 0 to 2 involve almost no technology spend and deliver a large share of the total available benefit, which is why the sequence matters so much. The durations assume sustained management attention; they stretch considerably without it. And stage 5 is where the real money is, because a failure prevented by a better pump specification never needs monitoring at all. Predictive maintenance is a stop on that road, not the end of it. The United States Department of Energy's federal energy management programme publishes a widely used operations and maintenance best practices resource that covers this progression for building and plant systems in useful detail.

11. Where predictive maintenance does not belong

The house rule: any strategy worth recommending is worth being honest about. Predictive maintenance is the wrong answer in several common situations.

  • Failure modes with no detectable warning. Electronic control boards, sudden fractures, and many instrument failures give effectively no advance signal. No amount of monitoring creates a warning period that physics does not provide. For these, the answer is redundancy, spares strategy, or a design change.
  • Low-consequence assets. If failure means a minor inconvenience and a cheap replacement, the monitoring cost cannot be recovered. Run to failure with a responsive corrective process is the economically correct strategy, and choosing it deliberately is a sign of maturity, not neglect.
  • Assets nearing end of life. Instrumenting equipment you intend to replace within two years rarely pays. Put the money into the replacement specification.
  • Statutory and safety-critical inspection. Where a regulation or an insurer requires an inspection at a defined interval, condition evidence does not release you from it. You can often reduce the intrusive work around it, but the mandated task stays.
  • Highly variable duty with no baseline of normal. Anomaly detection needs a stable definition of normal. An asset whose operating regime changes constantly and unpredictably is a poor candidate until the duty is better characterised.
  • Organisations without the prerequisites. The hardest one to say to a client who has already approved the budget. If the register, the history, the failure-mode understanding or the named owner is missing, a predictive programme will not work, and starting it anyway produces a failed initiative that makes the next attempt politically much harder.

A programme that covers fifteen percent of the register well is a success. One that covers everything thinly is a budget line with no reliability improvement attached, and it is far more common.

The idea to walk away with

Predictive maintenance is a decision-making capability, not a purchase. It works when evidence of a developing fault reliably becomes a planned intervention in the system where maintenance actually happens, on assets whose failure is both consequential and detectable, reviewed by somebody whose job it is to look. Everything else, the sensors, the platforms, the models, is infrastructure serving that loop.

Which means the honest sequence is the reverse of the usual one. Rank your assets by consequence. Clean your failure history. Understand how your equipment actually fails. Name the person who will act. Then, and only then, choose the measurement technology, starting with the data you already have and the cheapest techniques on the table above. Organisations that follow that order tend to end up with a programme that survives the second budget cycle. Organisations that buy the platform first and hunt for the use case afterwards tend to end up with the eighteen-months-later story, and an expensive dashboard nobody opens.

Final thoughts

The maintenance industry has spent twenty years describing predictive maintenance as a destination, and that framing has done real damage. It pushes organisations to skip the foundations, because the foundations look like going backwards, and it leaves them measuring progress by how much technology they have installed rather than by whether their assets fail less often.

The better framing is narrower and more useful. Predictive maintenance is one tool in a strategy mix, the most powerful and the most demanding, deployed where the evidence says it will pay. Most of the reliability improvement available to a typical operation sits in the stages before it: an accurate asset register, disciplined work order closeout, criticality-driven prioritisation, and a PM programme that has been reviewed for actual value. Those are unglamorous and they are where I would spend the first year of almost any engagement.

Do that groundwork, and predictive maintenance becomes a straightforward next step on a handful of assets where it obviously belongs. Skip it, and no sensor, platform or model will save the programme. The technology has not been the constraint for some years now. The organisational discipline always has been.

Considering a predictive maintenance programme?

Independent advisory on readiness assessment, criticality ranking, condition-monitoring strategy, CMMS and EAM integration, and the KPI framework to prove the programme worked. 22+ years across utilities, oil and gas, manufacturing, government and facility operations. No sensor vendor margins, no reseller arrangements.

Book a conversation

Related reading: Predictive maintenance and failure prediction, Preventive vs predictive vs reactive maintenance, The complete guide to preventive maintenance, Preventive maintenance strategies, Asset criticality classification, Reliability-centred maintenance, IoT integration with a CMMS, AI anomaly detection for early fault warning.

Muhammad Abbas

CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.

Work with me
MAbbaz.com
© MAbbaz.com