Most condition monitoring programmes I have reviewed do not fail because the wrong technology was purchased. They fail because the technique was pointed at a failure mode it physically cannot see, or because the measurement was taken in a way that made the number meaningless. A vibration reading from a hand-held probe pressed against a painted guard is not data. A thermal image of an unloaded circuit breaker is not evidence of anything. An oil sample drawn from the bottom of a drained sump tells you about the sump, not the gearbox. Technique selection and measurement discipline are the whole game, and they are both cheaper to get right than the instruments are to buy.
The message up front: each condition monitoring technique has a narrow band of failure modes it detects well and a much larger set it is blind to. Choose the technique from the dominant failure modes of the asset, not from the instrument you happen to own. Then fix the measurement location, the operating condition and the baseline, because a trend against your own baseline is worth far more than a single reading compared against a threshold copied from an article.
1. What condition monitoring is actually for
Condition monitoring exists to find the point on the degradation curve where a developing fault first becomes detectable, and to give you enough warning to intervene on your own schedule. It is the sensing layer underneath everything else. The analysis, the trending, the remaining-useful-life estimates and the machine learning all sit on top of it, and none of them can be better than the measurement they consume. If you are still placing condition monitoring in the wider strategy, the predictive maintenance practitioner's guide sets the frame, and the distinction that matters most for programme design is covered in condition-based versus predictive maintenance.
Three questions decide whether any technique belongs on a given asset:
- Does the dominant failure mode produce a detectable symptom? Bearing degradation produces high-frequency energy, heat and wear metal. A cracked control board produces nothing measurable until it stops working. No symptom, no technique.
- Does the technique see that specific symptom? Vibration sees rotating-element faults extremely well and sees insulation degradation not at all. Thermography sees resistive heating and sees a developing gear tooth crack only very late, if ever.
- Is the measurement interval shorter than the warning period? A fault that develops from detectable to functional failure in three weeks will not be caught by a quarterly route. This single mismatch accounts for a large share of the surprise failures inside programmes that technically have monitoring in place.
That third question is where route design lives, and it is worth noting that the answer differs by technique on the same machine. Ultrasound may detect a lubrication problem on a bearing months before vibration shows a defect frequency, which means the two techniques want different intervals even though they are watching the same component.
2. The master comparison: technique by technique
Before going deep on any one method, it helps to see the whole set side by side. The intervals below are common starting points for industrial rotating plant and electrical distribution, not prescriptions. Adjust them to criticality, duty and the warning period of the failure modes you are actually chasing. Asset criticality classification is the input that should drive the interval, not convenience of scheduling.
| Technique | Detects well | Cannot detect | Skill needed | Typical interval |
|---|---|---|---|---|
| Vibration analysis | Unbalance, misalignment, mechanical looseness, bearing defects, gear faults, blade and vane problems, cavitation, resonance, rotor-bar issues | Electrical insulation condition, lubricant chemistry, internal corrosion, faults on very slow or non-rotating assets, fluid contamination | High. Overall levels can be collected by a trained technician; spectral diagnosis needs a certified analyst | Monthly routes on critical rotating plant; continuous on high-consequence machines |
| Infrared thermography | Loose or high-resistance electrical connections, phase imbalance, overloaded circuits, failed capacitors, blocked cooling, insulation and refractory loss, steam trap and building envelope problems | Anything under load-free conditions, faults behind enclosures with no line of sight, internal mechanical wear, early bearing defects, anything where the heat signature is masked | Medium to high. Camera operation is easy; correct emissivity, reflected temperature and severity judgement are not | Annual or semi-annual electrical surveys; quarterly on critical switchgear and mechanical plant |
| Oil and lubricant analysis | Wear metal generation and its source, lubricant degradation and oxidation, viscosity change, water and coolant ingress, additive depletion, particle contamination, wrong oil in the machine | Faults that do not shed particles or alter chemistry: unbalance, misalignment, looseness, electrical faults, anything in a sealed or grease-lubricated component with no sample point | Low to collect if sampling is proceduralised; high to interpret, usually via an accredited laboratory | Quarterly on gearboxes and hydraulics; monthly on large or critical oil volumes; per oil change as a minimum |
| Ultrasound (airborne) | Compressed air, gas and vacuum leaks, steam trap condition, valve pass-through, corona, tracking and arcing in switchgear | Sealed internal mechanical wear, lubricant chemistry, anything requiring a spectrum to diagnose root cause, faults in noisy enclosures without access | Low to medium for leak surveys; medium to high for electrical discharge classification | Quarterly to semi-annual leak surveys; annual switchgear surveys alongside thermography |
| Ultrasound (structure-borne) | Very early bearing distress, under-lubrication and over-lubrication, cavitation, gear friction, bearing lubrication condition on slow-speed assets | Root cause of a rotating fault, unbalance and alignment problems, anything needing frequency-domain diagnosis | Low to medium. Decibel trending is accessible; time-waveform interpretation needs training | Monthly, often combined with the vibration route; weekly on bearings under active watch |
| Motor current signature analysis | Broken or cracked rotor bars, air-gap eccentricity, stator winding problems, supply imbalance, some driven-load faults, motor loading and efficiency | Mechanical faults on the driven machine that do not modulate current, lubricant condition, bearing defects in most practical cases | High for spectral interpretation; low for basic current and loading checks | Annual on critical motors; more often where rotor-bar history exists |
| Performance and thermodynamic monitoring | Efficiency loss, fouling, internal clearance wear, impeller and heat-exchanger degradation, filter loading, control problems | Component-level fault identification, anything that does not yet affect output, faults on assets without reliable instrumentation | Medium. Needs process understanding more than instrument skill | Continuous where instrumentation exists; otherwise monthly manual readings |
The pattern worth absorbing from that table: the techniques overlap far less than most programmes assume. Vibration and oil analysis on a gearbox are not redundant, they are complementary, because one sees the dynamics and the other sees the chemistry and the debris. Thermography and ultrasound on switchgear are likewise complementary, because a loose connection under load produces heat while a surface discharge produces sound long before it produces measurable heat. Buying one technique and assuming it covers the machine is the most common technical error in this space.
3. Vibration analysis: what the spectrum is telling you
Vibration is the workhorse for rotating equipment because almost every mechanical fault changes the way a machine moves, and it changes it at frequencies tied to the geometry and speed of the components involved. That is the core insight. A rotating machine generates vibration at running speed, at multiples of running speed, at tooth-passing and blade-passing frequencies, at bearing defect frequencies and at its own natural frequencies. When something develops, energy appears or grows at the frequency belonging to the component at fault.
There are two very different levels of practice, and confusing them is a common source of wasted money:
- Overall level monitoring gives you a single number, usually broadband velocity in millimetres per second RMS over a defined frequency range. It tells you the machine is getting worse. It does not tell you why. It is cheap, fast, easy to trend and perfectly adequate as a screening layer across a large population of non-critical machines.
- Spectral analysis transforms the time waveform into the frequency domain so you can see which components are contributing energy, and at what amplitude and phase. This is where diagnosis happens, and it needs a trained analyst plus accurate machine data: running speed, bearing part numbers, gear tooth counts, number of pump vanes or fan blades. Without that reference data you have a picture with no legend.
The time waveform matters too, and it is regularly skipped. Impacting, looseness and rubs often show more clearly in the waveform than in the spectrum, because the spectrum averages away the shape of individual events. A good analyst looks at both, plus phase between measurement points when diagnosing unbalance against misalignment.
The measurement discipline that makes trending possible
Same point, same direction, same mounting, same operating condition, every time. Mark the measurement locations permanently and fit stud-mounted or epoxy-mounted pads on critical machines. A magnet on a curved painted surface at a slightly different angle will change the high-frequency response enough to swamp the fault you are trying to trend. A perfect analyser with inconsistent measurement points produces a trend that is mostly measurement noise.
4. Common vibration fault signatures
The table below is the classic signature set that every vibration course teaches and every analyst carries in their head. Treat it as a diagnostic starting point, not a rule book: real machines produce mixed signatures, and structural resonance can amplify a small fault into an alarming amplitude or hide a serious one. The convention below writes running speed as 1X.
| Fault | Characteristic signature | Dominant direction | Confirming checks |
|---|---|---|---|
| Unbalance | Dominant 1X, amplitude rising with the square of speed, low harmonic content, stable phase | Radial, similar amplitude horizontal and vertical | Phase difference of roughly 90 degrees between horizontal and vertical at the same bearing |
| Misalignment (angular) | Strong 1X and 2X, sometimes 3X, high axial content | Axial, across the coupling | Roughly 180 degree axial phase shift across the coupling; check after any thermal growth has settled |
| Misalignment (parallel) | 2X often exceeding 1X, radial harmonics | Radial | Laser alignment check; confirm soft foot and pipe strain first |
| Mechanical looseness | Multiple harmonics of 1X, sometimes half-order harmonics, truncated or erratic time waveform | Radial, direction of the loose joint | Torque check on hold-down bolts, grout and baseplate inspection, soft foot test |
| Rolling element bearing defect | Non-synchronous defect frequencies (outer race, inner race, ball, cage) with sidebands, plus a rising high-frequency noise floor; envelope and demodulation techniques show it earliest | Radial, at the bearing housing | Calculate defect frequencies from bearing geometry; confirm with ultrasound and, where oil-lubricated, with wear particle analysis |
| Gear mesh problems | Gear mesh frequency (teeth multiplied by shaft speed) with sidebands spaced at shaft speed; sideband growth is the early indicator | Radial and axial depending on gear type | Borescope or strip inspection; oil analysis for gear wear metals |
| Blade or vane pass problems | Energy at blade or vane pass frequency and harmonics; often linked to clearance or flow issues | Radial and axial | Check clearances, inlet condition and system flow against the design curve |
| Cavitation | Random broadband high-frequency energy, no clear discrete peak, sounds like gravel | Radial at the pump | Suction conditions and net positive suction head available against required; ultrasound confirms readily |
| Resonance | One sharply amplified frequency that does not shift with speed changes; high amplitude in one direction only | Direction of the flexible mode | Bump test or coast-down; the fix is structural stiffening or speed change, not balancing |
| Electrical rotor problems | Pole-pass sidebands around 1X, twice line frequency content | Radial | Motor current signature analysis is the better confirming technique |
The important habit to build from this table is that the signature suggests and the confirming check decides. I have seen machines balanced repeatedly because the analyst stopped at a high 1X, when the actual problem was a resonance amplifying a modest unbalance, or a soft foot changing the stiffness in one direction. Balancing a resonant machine buys a few weeks. Confirming before acting is the difference between a diagnosis and a guess.
5. ISO velocity bands: how to use them, and how not to
The ISO 10816 series, now largely superseded by the ISO 20816 series, gives broadband vibration evaluation criteria for machines grouped into classes by size, power, mounting and type. It divides measured velocity into zones that are conventionally described as good for a new machine, acceptable for unrestricted long-term operation, unsatisfactory, and unacceptable. Those zone boundaries are the numbers people copy into alarm limits.
Copying them without reading the standard is a mistake, for several reasons:
- The boundaries depend on machine class. A small directly-coupled pump and a large turbine-generator on a flexible foundation have very different limits. Using one table for the whole plant will over-alarm some machines and under-alarm the ones that matter.
- Mounting matters. Rigid and flexible support classifications give different limits for otherwise identical machines, and misclassifying the mounting shifts every number.
- The measurement definition matters. The criteria apply to a defined quantity over a defined frequency range at defined locations. A reading taken over a different frequency band is not comparable to the table value.
- Revisions change. Class definitions and numbering have moved between the 10816 and 20816 series. Quoting a figure from an old summary can mean quoting a limit that no longer applies to that machine group.
- Standards evaluate acceptability, not health trends. A machine sitting comfortably inside the acceptable zone but with vibration that has tripled in two months is a machine with a developing fault, and no absolute limit will tell you that.
Do not take a threshold from an article, including this one
I have deliberately not reproduced the zone boundary values here. Buy the current standard from ISO , classify each machine group properly, and then set your own alarm limits from the standard plus your own baseline data. The honest position is that your baseline is more useful than any published limit: establish what each machine reads when it is known healthy at normal duty, and alarm on percentage change from that baseline as well as on the absolute zone. The absolute limit catches a machine that was always bad. The baseline trend catches the machine that is going bad, which is the one you actually want to find.
6. Route-based collection against continuous monitoring
Both models are legitimate, and the choice should be driven by the warning period of the failure modes involved and the consequence of missing one.
Route-based collection means a technician walks a defined route with a data collector on a fixed cycle, typically monthly. It is cost-effective, it spreads one analyst across hundreds of machines, and the human presence has real diagnostic value: the technician hears, feels and smells things that no sensor reports. Its weakness is interval. A fault that goes from detectable to failed inside the route cycle will be missed, and the data is a set of snapshots taken under whatever operating condition happened to prevail on the day.
Continuous monitoring, whether wired systems on large machines or wireless sensors on smaller ones, removes the interval problem and captures transient and load-dependent behaviour that a monthly snapshot cannot. It also removes the technician from the machine, which means you lose the walk-around observations, and it introduces a new cost centre in data volume, connectivity, battery or power management and alarm administration. The architecture considerations are covered in IoT sensors for predictive maintenance.
The pattern I would recommend for most industrial and facilities estates is a tiered one:
Tier 2, important with partial redundancy → monthly route with full spectral collection, plus overall-level wireless sensors where access is difficult or unsafe.
Tier 3, low consequence or easily replaced → overall level screening on a longer cycle, or no vibration monitoring at all, with an appropriate preventive or run-to-failure strategy instead.
That tiering is a criticality decision before it is a technology decision. Putting continuous monitoring on a machine with a standby unit and a two-hour swap time, while the single unredundanted critical pump sits on a quarterly route, is a resource allocation failure that no amount of analytics fixes. The strategy split behind this is explored in preventive versus predictive versus reactive maintenance.
7. Thermography: the two things people get wrong
Infrared thermography is the most accessible of the four main techniques and, partly for that reason, the most frequently done badly. A modern camera produces an attractive image with a temperature readout on almost any surface, and it is very easy to believe the number. The number is a calculation, and two inputs drive it.
Emissivity is the efficiency with which a surface radiates infrared energy. A painted or oxidised surface has high emissivity and behaves predictably. Bare or polished metal, which is exactly what busbars, lugs and copper connections are made of, has low emissivity and will read dramatically colder than it actually is. A shiny connection that is genuinely running hot can appear cool on the screen. The practical workarounds are to target adjacent high-emissivity surfaces such as insulation or paint, to apply emissivity targets or tape at fixed inspection points where it is safe to do so, and to set the camera emissivity to match the surface rather than leaving it at a default.
Reflected temperature is the infrared energy from the surroundings that bounces off the target into the camera. A low-emissivity surface is a good mirror, which means a polished panel can show you the reflection of a hot lamp, a warm body or direct sun and present it as target temperature. Change the viewing angle, and if the hot spot moves across the surface it is a reflection. If it stays put, it is real. That single test resolves most false positives.
Beyond those two, the recurring errors are practical:
- Surveying an unloaded circuit. Resistive heating is proportional to the square of current. At low load a serious loose connection produces almost no detectable rise. An electrical thermography survey done during a shutdown, or on a standby feeder, or at night on a daytime load, tells you essentially nothing. Load condition must be recorded with every image, and surveys should be scheduled at representative load.
- Reporting absolute temperature instead of delta-T. A 60 degree busbar in a 45 degree switchroom in an Abu Dhabi summer is a different proposition from a 60 degree busbar in a 20 degree plant room. Severity criteria are normally expressed as a temperature rise: against the same component on the other phases under similar load, against a similar component elsewhere in the same installation, or against ambient. Phase-to-phase comparison at similar loading is usually the most defensible.
- Treating published delta-T severity bands as absolute rules. Commonly used severity frameworks group findings into tiers such as monitor, repair at next opportunity, repair soon and immediate action, with delta-T thresholds attached. Those thresholds vary between publications, between the phase-comparison and ambient-comparison bases, and between equipment types. Work from a recognised electrical maintenance testing framework, for instance the specifications published by NETA , and record which basis and which criteria you used in the report so the next person can compare like with like.
- Missing the line of sight. Infrared does not see through enclosure doors, glass or most plastics. A survey with panels closed is a survey of panel doors. Either use installed infrared windows or plan for safe opening under a proper electrical safety procedure, which is a significant scheduling and permit consideration in itself.
Thermography is not only an electrical technique. On the mechanical side it finds blocked heat exchangers and radiators, failed steam traps, refractory and insulation loss, hot bearings and couplings, motor cooling problems, and moisture in building fabric. The caveat is that mechanical thermography is usually a late indicator: by the time a bearing is measurably hot at the housing, vibration and ultrasound have typically had the fault for some time. For the routine checks around electrical assets, the electrical preventive maintenance checklists give the surrounding task set, and thermal survey scheduling on thermal plant fits into the boiler and chiller preventive maintenance routines.
8. Oil analysis: what the sample can and cannot tell you
Oil analysis is the technique with the longest warning period on lubricated machinery and the one most often reduced to a filing exercise, where samples are sent, reports arrive, and nobody reads them unless a value is flagged red. Used properly it answers three separate questions, and it is worth keeping them separate because they call for different tests.
Question one: is the machine wearing? This is wear debris analysis. Spectrometric analysis, commonly by inductively coupled plasma or rotating disc electrode, measures the concentration of wear metals, contaminant elements and additive elements in parts per million. Iron points to gears, shafts and bearing races. Copper suggests bushings, thrust washers or a cooler. Chromium and nickel indicate alloy steel components. Aluminium suggests pistons, housings or certain bearings. Tin and lead point to bearing overlays. Silicon usually means dirt ingress, occasionally an antifoam additive or a sealant.
The critical limitation of spectrometric analysis is particle size: the technique responds well to small particles and progressively underreports larger ones, with the practical ceiling for routine methods being in the low single-digit to roughly ten micron region depending on method. Advanced wear, which generates large particles, can therefore show a flat or even falling trend on the elemental report while the machine is failing. That is why particle quantifier indices, ferrous density measurement and analytical ferrography, where particles are examined under a microscope and classified by morphology as cutting, sliding, fatigue or corrosive wear, are added on critical machines. Ferrography also tells you the wear mechanism, not just the quantity, which is what actually guides the intervention.
Question two: is the oil still fit for service? Viscosity at 40 and 100 degrees Celsius is the single most important lubricant property, because viscosity is what generates the film that separates surfaces. A rise usually means oxidation, contamination with a heavier fluid, or soot loading. A fall usually means fuel dilution, contamination with a lighter fluid, shear of the viscosity modifier, or simply the wrong oil having been added. Oxidation and nitration by infrared spectroscopy track chemical degradation. Acid number rising indicates oxidation and acidic by-product formation in industrial oils; base number falling indicates depleted alkaline reserve in engine oils. Additive elements such as zinc, phosphorus, calcium and magnesium trending downward indicate additive depletion, and a change in the additive fingerprint is often the first sign someone topped up with a different product.
Question three: is the oil contaminated? Water is the contaminant that does the most damage for the least visible presence, measured by crackle test for screening and Karl Fischer titration for quantification, and even modest dissolved water concentrations reduce bearing life meaningfully. Glycol indicates a cooler or head gasket leak. Fuel dilution indicates injection or combustion problems. Particle counting reports the number of particles above set size thresholds and is expressed as an ISO 4406 cleanliness code, a set of range numbers for particles greater than 4, 6 and 14 microns. Hydraulic and turbine systems live or die by this number, and component manufacturers publish target codes for their equipment. Take the target from the component manufacturer and the coding convention from the current standard at ISO , because the code definition has itself been revised.
Sampling discipline decides whether oil analysis works at all
Sample from a permanently fitted, live-zone sampling valve upstream of the filter and downstream of the component, while the machine is running at normal operating temperature. Flush the valve and the tubing before taking the sample. Use clean, certified bottles and never decant. Sample the same point, the same way, at the same interval, every time. A sample drawn from a drain plug during a drain gives you the settled sludge of the whole service interval, not the current condition of the machine, and it will produce a trend you cannot interpret. If I had to choose between a good laboratory with sloppy sampling and an average laboratory with disciplined sampling, I would take the second every time.
The other point about oil analysis is that a single report is nearly worthless. The value is in the trend against the machine's own history, because acceptable wear metal levels vary hugely by machine type, oil volume, filtration and service life. Laboratory alarm limits are generic starting points. Your own trend, with oil changes and top-ups annotated so you can see when the baseline was reset, is the real instrument. Oil volume matters here too: the same absolute mass of wear debris produces a much higher concentration in a small gearbox than in a large circulating system, so concentration trends are only comparable within the same machine.
9. Ultrasound: often the earliest warning you can get
Ultrasound detects sound above the range of human hearing, typically monitored in a band around 20 to 100 kilohertz, and heterodyned down into the audible range so the operator can listen to it. It splits into two applications that share an instrument but not much else.
Airborne ultrasound is for anything that leaks or discharges into the air. Turbulent flow through a leak generates broadband ultrasound, so compressed air, gas and vacuum leaks are found quickly and located precisely, even in a noisy plant, because the high-frequency energy is strongly directional and attenuates fast. Compressed air leak surveys are one of the few condition monitoring activities with a return that finance functions accept without argument, since the energy cost of leakage is directly calculable. The same technique classifies steam trap condition, finds valve pass-through, and detects corona, surface tracking and arcing in switchgear. That last application is important precisely because electrical discharge is often ultrasonically loud long before it is thermally hot, which makes ultrasound and thermography a genuinely complementary switchgear pair rather than duplicate spending.
Structure-borne ultrasound uses a contact probe on a bearing housing or casing. Here the value is early bearing and lubrication condition. As a rolling element bearing begins to lose its lubricant film, metal-to-metal contact raises the ultrasonic amplitude before any discrete defect frequency appears in the vibration spectrum, because there is no defect yet, only friction. That is why structure-borne ultrasound frequently detects earlier than vibration: it responds to the friction and micro-impacting that precedes the formation of a measurable fault. The practical payoff is lubrication done on condition rather than on calendar, which also addresses over-greasing, a genuinely common cause of bearing failure that a pure schedule-based regime actively creates. It is also the technique of choice for slow-speed bearings, where vibration analysis struggles because there is very little energy at low rotational speeds.
What ultrasound cannot give you
Ultrasound tells you something is wrong and roughly how wrong, and it tells you early. It does not tell you what is wrong in the way a vibration spectrum does. It will not distinguish unbalance from misalignment, it will not identify which bearing race has spalled, and decibel readings are strongly dependent on probe pressure, contact point and instrument settings, so inconsistent technique destroys the trend even faster than it does in vibration work. Use ultrasound as the early detector and the lubrication tool, and use vibration as the diagnostic that tells you what to do about it.
10. Motor current signature analysis and performance monitoring
Two further techniques deserve a place in the toolkit because they cover gaps the main four leave open.
Motor current signature analysis treats the motor as its own transducer. Mechanical asymmetries inside the motor, and some in the driven load, modulate the stator current, so a spectrum of the current signal reveals broken or cracked rotor bars, high-resistance end ring joints, air-gap eccentricity, stator winding problems and supply imbalance. Its two practical advantages are that measurement is taken at the motor control centre rather than at the machine, which means no access to a hazardous or remote location and no shutdown, and that it sees rotor-cage faults which vibration detects poorly. Its limitations are real: it needs the motor loaded to be meaningful, spectral interpretation is a specialist skill, and it is weak on the mechanical faults of the driven equipment. Treat it as the electrical complement to vibration on critical motors, not a replacement. Offline electrical testing, insulation resistance, polarisation index and winding resistance, sits alongside it and answers different questions about insulation condition that no online technique addresses well.
Performance and thermodynamic monitoring is the most underused technique on this list and usually the cheapest, because the instruments are already installed. Flow, pressure, temperature, speed, current and power readings from the BMS, SCADA or plant historian encode condition information. A pump whose differential head is falling at constant speed and flow is wearing internally. A chiller whose approach temperatures are widening is fouling. A heat exchanger whose effectiveness is dropping is scaling. A filter whose differential pressure is climbing is loading. None of this needs a new sensor, only someone to define the expected performance envelope and trend against it. For the asset classes where this pays best, see generator, pump and motor preventive maintenance.
The limitation of performance monitoring is that it identifies degradation without identifying the component. It tells you the pump is 8 percent less efficient than its baseline; it does not tell you whether that is impeller wear, a clearance problem or a partially blocked strainer. It is an excellent screening layer that hands off to a diagnostic technique, and it is a poor diagnostic technique on its own.
11. The honest cost and skill picture
This is the part that gets skipped in vendor presentations, and it is the part that determines whether a programme survives its third year. I will give the shape of the cost rather than invented figures, because equipment pricing, laboratory rates and labour costs vary too much by region and contract for a published number to be useful.
- Vibration analysis carries the highest total cost of ownership. A route-capable analyser with spectral capability, software licensing, accelerometers and mounting hardware is a significant capital item, and it is dwarfed over time by the analyst. Certified analyst capability is a multi-year development path with recognised certification levels, and a competent analyst is an expensive, mobile and hard-to-replace resource. Outsourcing the analysis while keeping the collection in-house is a reasonable middle path, but be aware that a contractor collecting quarterly is often selling you a compliance report rather than a reliability outcome.
- Thermography has the lowest barrier to entry and a deceptively low apparent cost. Cameras have become affordable, and the temptation is to hand one to a technician and call it a programme. The cost that matters is training to a recognised certification level, the time to plan surveys at representative load, the electrical safety arrangements to get line of sight into energised panels, and the reporting discipline to make findings actionable. The camera is the cheap part.
- Oil analysis has the lowest capital cost of all: sampling valves, bottles, pumps and a laboratory contract. Per-sample laboratory costs are modest, which is why this is usually the best value first technique on lubricated machinery. The hidden cost is retrofitting proper sampling points to machines that do not have them, and the recurring cost is the engineering time to read and act on reports, which is exactly the cost most organisations decline to fund. Unread reports are pure waste.
- Ultrasound sits in the middle on equipment and low on training for leak work, which makes a compressed air leak programme one of the most reliably self-funding activities available. Bearing and lubrication work needs more training and, more importantly, needs procedural consistency to produce trendable numbers.
- Motor current signature analysis is specialist equipment and specialist interpretation, most often bought in as a service on a defined population of critical motors rather than built in-house.
Where condition monitoring does not pay
On low-consequence, low-cost, easily replaced assets, condition monitoring costs more than the failures it prevents. On assets whose dominant failure mode is sudden, monitoring produces data and no warning. On assets with full redundancy and fast changeover, the consequence being avoided may be too small to fund the technique. And in an organisation with no capacity to act on findings, every technique on this list is a cost with no return: a programme that generates reports nobody converts into scheduled work is worse than no programme, because it consumes budget and manufactures false confidence. Be willing to say no to monitoring. The discipline to leave most of the asset register unmonitored is what makes it affordable to monitor the few assets properly.
12. Building a programme that produces decisions
The technical content above is necessary and not sufficient. A working programme needs a small number of structural things in place, and in my experience these are what separate the programmes that last from the ones that quietly stop.
- Start from failure modes, not from instruments. For each critical asset class, list the dominant failure modes, then assign the technique that detects each one. Where a failure mode has no detectable symptom, record that honestly and address it with redundancy or design change instead.
- Fix the measurement points before the first reading. Permanent vibration pads, permanent sampling valves, infrared windows, labelled ultrasound contact points. Retrofitting these later invalidates the baseline you have already collected.
- Record the operating condition with every reading. Load, speed, flow, ambient temperature. A reading without its operating context cannot be compared to another reading, and this is the most common reason a historical database turns out to be unusable.
- Build the baseline deliberately. Take readings when the machine is known good, ideally after commissioning or overhaul, and treat that as the reference. Reset and annotate the baseline whenever the machine is rebuilt or the oil is changed.
- Alarm on change as well as on absolute value. Absolute limits from a standard catch machines that are already bad. Rate-of-change alarms against baseline catch machines that are becoming bad, which is the point of the whole exercise.
- Route findings into the CMMS as work, not as a report. A finding that lands as an email or a PDF attachment has no owner, no due date and no closure record. It needs to become a work order in the same system where the technician already works, with the recommendation, the severity and the evidence attached. Platforms such as IBM Maximo, Hexagon EAM, Infor EAM, SAP PM and Planon all support condition-driven work generation, as do the lighter tools like Fiix, eMaint, Limble, MaintainX and UpKeep, but the integration has to be configured deliberately rather than assumed.
- Close the loop with findings verification. When a machine is opened, record what was actually found against what the technique predicted. Without that feedback, nobody learns whether the diagnosis was right, severity criteria never get tuned to your plant, and the programme cannot demonstrate its own value.
That last point is the one I would push hardest. A condition monitoring programme that cannot show a list of faults it found early, and what those findings avoided, will lose its budget in the first cost review. The technical work is only half the job; documenting the outcome is what keeps it funded. The prediction layer that sits above all of this, including where machine learning genuinely helps, is covered in predictive maintenance and failure prediction.
The idea to walk away with
Each of these techniques is a narrow instrument with a specific sensitivity. Vibration sees mechanical dynamics. Thermography sees resistive heating and thermal anomalies with line of sight. Oil analysis sees wear chemistry and debris in lubricated systems. Ultrasound sees friction, turbulence and discharge, usually earliest of all. Motor current analysis sees electrical asymmetry. Performance monitoring sees the consequences of degradation without naming the cause. Selecting from that set by failure mode, and combining techniques where a machine has several failure modes that no single technique covers, is what a technically competent programme looks like.
And underneath the technique selection sits the measurement discipline, which is unglamorous and decisive. Fixed points, consistent mounting, representative load, controlled sampling, recorded operating context, and a baseline you built yourself. Get those right with mid-range instruments and you will find faults. Get them wrong with the best instruments available and you will generate a large, expensive database of numbers that cannot be compared to each other.
Final thoughts
If you are starting from nothing, the sequence I would advise is not glamorous. Begin with oil analysis on your lubricated critical machinery, because the capital cost is low and the warning period is long. Add an ultrasound leak survey, because it tends to pay for the instrument out of energy savings and builds credibility for the programme with the people who control the budget. Add thermography on electrical distribution, done properly, at representative load, with emissivity and reflected temperature understood. Then build vibration capability on the critical rotating population, either in-house with a proper analyst development path or through a contract where you own the data and the measurement points.
And whatever you do, take your thresholds from the current standards and your own baselines, not from a table on a website. The numbers in the standards are tied to machine class, mounting, measurement definition and revision, and the numbers that actually catch your developing faults are the ones you derive from your own healthy-machine readings. That is the difference between running a condition monitoring programme and owning an instrument.
Designing or reviewing a condition monitoring programme?
Independent advisory on technique selection by failure mode, route and interval design, measurement point standards, and getting condition findings into the CMMS as real, closed-out work. 22+ years across utilities, oil and gas, manufacturing, government and facility operations. No instrument vendor margins, no reseller arrangements.
Book a conversationRelated reading: Predictive maintenance: a practitioner's guide, Condition-based vs predictive maintenance, Predictive maintenance and failure prediction, IoT sensors for predictive maintenance, Electrical preventive maintenance checklists, Asset criticality classification.
Muhammad Abbas
CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.
Work with me