Ask ten maintenance engineers for a list of failure modes on a centrifugal pump and you will get ten lists, most of them mixing up four completely different kinds of statement. Some entries will be functional failures, some will be mechanisms, some will be causes, and one or two will simply be the name of a component. That confusion is not pedantry, it has consequences: a failure-mode list built on muddled definitions cannot support a criticality decision, cannot drive a maintenance task, and cannot be turned into a usable code set. So this guide starts with the definitional discipline, because that is where the real-world lists go wrong, and only then moves on to the taxonomy of common types.
The message up front: a failure mode names the item and the specific manner in which it fails, and a well-written one points straight at a checkable physical mechanism. "Pump failed" is not a failure mode. Neither is "bearing failure". If your list is full of entries like those, no analysis built on top of it will be worth much, however sophisticated the method.
1. The four levels: functional failure, mode, mechanism, cause
The single most useful thing you can do to a failure-mode list is separate four levels of statement that people routinely collapse into one. Each answers a different question, and each is used for a different purpose downstream.
- Functional failure: the required function is no longer delivered to the required standard. This is defined against a performance requirement, not against a component. It is the level at which operations feels the problem.
- Failure mode: the specific manner in which the item fails. It names an item and a manner of failing. This is the level at which maintenance decisions are made.
- Failure mechanism: the physical, chemical or electrical process that produced the mode. Fatigue, abrasion, pitting corrosion, insulation thermal ageing. This is the level at which engineering explains the mode.
- Failure cause: why that mechanism was active in the first place. Wrong lubricant specified, misalignment at installation, cooling water chemistry outside limits, a protective setting left at default. This is the level at which the problem is actually eliminated.
The clearest way to make the distinction stick is to carry one example all the way down. Take a chilled-water pump whose stated function is to deliver a specified flow at a specified head, continuously, during cooling season.
| Level | What it answers | Worked example carried through |
|---|---|---|
| Function | What is the item required to do? | Deliver 120 litres per second at 35 metres head, continuously, to the chilled-water ring main. (Illustrative figures.) |
| Functional failure | In what way has the required performance been lost? | Unable to deliver the required flow at the required head. Note this can be true while the pump is still turning. |
| Failure mode | In what specific manner did the item fail? | Drive-end bearing seizes due to loss of lubrication. Another distinct mode on the same functional failure: impeller wears so clearances open beyond limits. |
| Failure mechanism | What physical process caused the mode? | Adhesive wear and thermal seizure of the rolling elements once the lubricant film collapsed. |
| Failure cause | Why was that mechanism active? | Greasing task removed from the schedule during a PM rationalisation exercise, so the bearing had not been regreased in two years. |
| Effect / consequence | What happens as a result, and who cares? | Loss of cooling redundancy, elevated supply temperature if the standby pump is also unavailable, motor overload trip, secondary shaft and seal damage. |
Levels are illustrative and simplified; the point is that each row answers a different question and none substitutes for another.
Read that table from the bottom up and you get the anatomy of a root cause investigation. Read it from the top down and you get the anatomy of a design analysis. Both directions rely on the levels being kept separate. The moment "loss of lubrication" is recorded as the failure mode and the bearing seizure is forgotten, you have lost the ability to ask whether any other mode shares the same cause, or whether the same mode can arrive by a different route.
The practical test
If a statement tells you nothing about what to inspect, it is a functional failure and not a mode. If it tells you what to inspect but not what to correct permanently, it is a mode. If it tells you what to change in the design, the specification or the schedule, you have reached the cause. Three different outputs, three different levels.
2. How to write a failure mode that is actually usable
Here is the rule I apply in every workshop: a usable failure mode names the item and the manner, and the test of a well-written one is whether it points at a specific, checkable mechanism. If two engineers reading the same entry would go and look at different things, it is not written well enough.
That rule quietly disqualifies two of the most common entries in real registers. "Pump failed" is not a failure mode, it is a functional failure at best and an incident note at worst. And "bearing failure" is not a failure mode either, which surprises people, because it names an item but not a manner. A bearing can seize, it can spall, it can suffer cage fracture, it can lose preload, it can fail from electrical erosion under a variable-speed drive. Those have different warning signatures, different detection routes and different correct responses. Collapsing them into "bearing failure" throws away exactly the information that would have made the entry useful.
A short before-and-after list is the fastest way to communicate the standard. This table alone tends to change how a team writes for the rest of the exercise.
| Badly written | What is wrong | Well written | What it now enables |
|---|---|---|---|
| Pump failed | Functional failure, not a mode | Mechanical seal leaks past the primary faces, losing pumped fluid to the gland area | Seal leak inspection on round; seal upgrade decision |
| Bearing failure | Item named, manner missing | Non-drive-end bearing spalls on the outer race under fatigue loading | Vibration monitoring at bearing defect frequencies |
| Motor burnt out | Effect stated as mode | Stator winding insulation breaks down phase to earth | Insulation resistance and polarisation index testing |
| Poor maintenance | This is a cause, and a vague one | Coupling loosens progressively, allowing angular misalignment beyond tolerance | Laser alignment check; torque verification task |
| Corrosion | Mechanism, not a mode | Pipe wall thins by internal erosion corrosion downstream of the elbow until it perforates | Ultrasonic wall thickness survey at defined points |
| AHU not cooling | Symptom reported by the user | Chilled water control valve sticks in the partially open position | Valve stroke test; actuator replacement decision |
| Sensor problem | No item, no manner | Differential pressure transmitter drifts out of calibration, reading low | Scheduled calibration interval with recorded as-found values |
| Human error | Blame, not engineering | Isolation valve left closed after intervention, starving the pump suction | Return-to-service checklist and valve line-up verification |
Notice that every well-written entry on the right hand side implies a specific check. That is not a coincidence, it is the whole purpose. A failure mode written to this standard is a bridge between the engineering reality and the maintenance task. Written to the standard on the left, it is a label that survives the workshop and helps nobody afterwards.
One more discipline worth imposing: write the mode at a consistent level of the asset hierarchy. A list that mixes "pump set fails to start" with "third-stage impeller vane cracks" is not internally comparable, and any ranking built on it will be distorted by the level of detail rather than by real risk. Decide the analysis boundary first, usually the maintainable item, and hold the whole list to it. If you are still fixing your hierarchy, the asset hierarchy design guide covers the levels this depends on.
3. One item, several failure modes
A frequent objection in workshops is that one component should have one failure mode, and that multiple entries for the same item are duplication. They are not. A single item legitimately has several failure modes, each with a different mechanism, a different cause, a different rate of development and a different correct response. Reducing them to one entry is not tidying up, it is deleting information.
Take a motor-driven fan in an air handling unit. The belt drive alone can fail by belt wear and elongation reducing tension, by belt fracture from fatigue after ageing and cracking, by sheave wear widening the groove profile, and by slip from contamination with oil or water. Four modes, one subassembly. The wear mode develops over months and is detectable on a belt tension check. The fracture mode may be sudden and is better managed by age-based replacement. The contamination mode is fixed by finding the leak, not by touching the belt at all. If you collapsed all four into "belt failure", you would be forced to choose one response for four different problems.
The same logic applies in the other direction. Two different items can produce the same functional failure, and the analysis needs both. A fan that fails to deliver airflow can be a belt problem, a blocked filter, a closed damper, a failed motor, a reversed rotation after a phase swap, or a control system that never issued the start command. Any of those satisfies "no airflow". Only one of them will be the mode you actually face, which is why the functional failure is the wrong level to plan work from.
Where this gets expensive
Splitting modes finely is correct in principle and unaffordable without limit. Every mode you add is another row to review, rank and maintain. The discipline is to split only where the split changes the decision. If two modes share the same detection method, the same task and the same consequence, carrying them as one entry costs you nothing. If they differ on any of the three, keep them apart.
4. Hidden versus evident failure modes
This is one of the most consequential distinctions in the whole subject, and the one most often missing from failure-mode lists entirely. A failure mode is evident if its occurrence becomes apparent to the operating crew in the normal course of their duties. It is hidden if it does not: nothing announces it, no alarm sounds, nothing changes in the control room, and the organisation carries on unaware.
The reason this matters is uncomfortable once you see it. A hidden failure does not cause a problem on its own. It sits there, silently, until a second event calls on the failed item. The failed standby pump is discovered when the duty pump trips. The stuck relief valve is discovered when the pressure excursion arrives. The fire pump that will not start is discovered by the fire. The exposure is not the single failure, it is the coincidence of the hidden failure with the demand, and that exposure grows for every day the hidden failure goes undetected.
The equipment where hidden modes concentrate is predictable, and it is the equipment that exists precisely because something else might fail:
- Protective devices: relief valves, trip systems, earth fault protection, interlocks, emergency shutdown logic. Their function is to act on demand, and a failed one looks exactly like a healthy one.
- Standby and redundant equipment: standby pumps, N+1 chillers, backup generators, redundant power supplies, secondary network paths. A standby unit that will not start is invisible until the day the duty unit stops.
- Alarms, detection and annunciation: gas detection, smoke detection, level alarms, dial-out and paging. An alarm that has silently lost its output path is worse than no alarm, because it is trusted.
- Passive and on-demand features: fire dampers, pressure relief paths, drainage routes, battery capacity behind an uninterruptible supply.
The right response to a hidden failure mode is neither preventive maintenance in the usual sense nor a reactive repair, because you cannot react to something you have not noticed. The correct response is a failure-finding task: a deliberate test whose only purpose is to discover whether the hidden failure has already occurred. Starting the standby pump on a rotation, function-testing the trip circuit, stroking the damper, lifting the relief valve on a bench, injecting a test signal into the detection loop. The task does not prevent the failure and does not fix it. It converts a hidden failure into an evident one, which is the only thing that lets you manage the exposure at all.
This is exactly the ground that reliability centred maintenance is built on, and the evident-versus-hidden split is the first branch in its consequence logic. SAE JA1011_202411, "Evaluation Criteria for Reliability-Centered Maintenance (RCM) Processes", sets the criteria a process must satisfy before it may legitimately be called RCM, including how failure consequences are categorised and how tasks are selected. It sets criteria rather than prescribing a process, which is a distinction worth remembering when a vendor claims their tool "is" RCM. SAE JA1012_201108, "A Guide to the RCM Standard", is the accompanying guide, and is now one generation behind JA1011. For the method itself, see the RCM introduction, which covers the seven questions and the task selection logic in full.
The question to ask of every mode
Would the operating crew know, in the normal course of their work, that this had happened? If the honest answer is no, you are looking at a hidden failure mode and it needs a failure-finding task with a defined interval, defined pass criteria and a recorded as-found result. "We test it when we get a chance" is not a failure-finding task.
5. Degraded versus complete failure
The next distinction to get right is between an item that has stopped and an item that is no longer good enough. Both are failures, and treating only the first as real is how organisations end up with equipment that has been quietly out of specification for years.
A complete failure is loss of the function altogether: the pump does not turn, the transmitter outputs nothing, the valve will not move. A degraded or partial failure is loss of the required performance while the item continues to operate: the pump turns but no longer makes head, the transmitter reads but is drifting, the valve strokes but only to seventy percent of travel. A functional failure can therefore occur while the item is still running, and this is the point most maintenance systems handle badly, because a running asset with no work order looks healthy in every report.
Whether a degraded state counts as a failure depends entirely on how precisely you wrote the function. "Deliver chilled water" is satisfied by a pump limping along at half flow. "Deliver 120 litres per second at 35 metres head" is not. This is why the definition of the function has to carry a performance standard, and why the functional failure is defined against that standard rather than against the component. Vague functions generate lists that miss all the degraded modes, which are frequently the expensive ones because they run for a long time before anyone declares them broken.
Degraded modes are also the modes that make condition monitoring worthwhile. A mode that develops progressively passes through a detectable stage before it reaches functional failure, which is the interval the P-F curve describes and which the condition monitoring techniques guide gets its usefulness from. Modes that go from healthy to complete failure with no intermediate stage have no such interval, and no amount of monitoring will find them.
6. Common failure modes by mechanism
With the definitions settled, the taxonomy becomes genuinely useful. Organising by mechanism is the version engineers find most natural, because it groups modes by the physics that produces them, which in turn tells you what evidence to look for and whether the mode can be detected while the item is still in service. The evidence column here is deliberately brief; identifying a mechanism from the state of a failed part is the subject of the failure analysis methods guide, which covers fracture surfaces, wear evidence and when to escalate to a laboratory.
| Mechanism | What it is | Where it shows up | Typical evidence left behind | Detectable in service? |
|---|---|---|---|---|
| Mechanical wear and abrasion | Progressive material loss where surfaces move against each other or against particles | Impellers, wear rings, gears, bushes, chain and belt drives, guides | Polished or scored surfaces, directional scratches, dimensional loss, wear debris in oil | Often yes. Clearance checks, performance drift, wear particles in oil analysis |
| Fatigue | Crack initiation and growth under repeated or fluctuating stress, below the static strength | Shafts, bearing races, springs, welds, pipe supports, fan blades, bolting | Progressive crack markings then a final fast fracture zone; often initiates at a stress concentration | Sometimes. Vibration change, crack detection by NDT at known initiation sites |
| Overload and overstress | A single event exceeding design capacity, mechanical, electrical or hydraulic | Couplings, gear teeth, structural members, breakers, cables, lifting gear | Gross deformation, tearing, single-event fracture with no progressive markings | Rarely. Usually managed by protection, design margin and duty control |
| Corrosion and erosion | Electrochemical attack, or mechanical removal by fluid and entrained solids | Pipework, vessels, tube bundles, valve seats, structural steel, enclosures, earthing | Pitting, general wall loss, localised thinning at bends and downstream of restrictions, deposits | Yes. Wall thickness surveys, visual inspection, water chemistry trending |
| Lubrication related | Loss, degradation, contamination or simply the wrong lubricant, letting surfaces contact | Bearings, gearboxes, hydraulics, compressors, engine assemblies | Discolouration and heat tinting, smearing, degraded or contaminated oil, seized elements | Yes, and usually early. Oil analysis, temperature, ultrasound at the bearing |
| Thermal and overheating | Operation above design temperature, or loss of cooling, degrading materials | Motors, transformers, VFDs, engine cooling circuits, insulation, elastomers | Discolouration, embrittled or hardened polymers, varnish, distortion, blocked cooling paths | Yes. Thermography, temperature trending, cooling differential checks |
| Electrical | Insulation breakdown, contact degradation, winding faults, partial discharge, tracking | Switchgear, motors, transformers, cabling, terminations, control panels | Carbon tracking, pitted or eroded contacts, charred insulation, arcing marks | Yes for several. Insulation resistance, thermography, partial discharge, contact resistance |
| Contamination and blockage | Foreign material obstructing a flow path or reaching a clearance it should not | Filters, strainers, coils, nozzles, drains, heat exchangers, hydraulic circuits | Fouling deposits, clogged media, differential pressure rise, reduced flow | Yes, and cheaply. Differential pressure, flow, visual inspection |
| Loosening, misalignment, imbalance | Loss of the intended geometric relationship between parts | Couplings, baseplates, fan and pump rotors, bolted joints, drive trains | Fretting at joint faces, elongated holes, uneven wear patterns, witness marks | Yes. Vibration analysis is strong here; alignment and torque verification |
| Sealing and leakage | Loss of containment at a static or dynamic seal, gasket or joint | Mechanical seals, gland packing, flanges, O-rings, hydraulic cylinders, refrigerant circuits | Staining and drip evidence, hardened or extruded elastomers, scored seal faces | Usually yes. Visual inspection, pressure decay, ultrasonic and tracer leak detection |
| Control, instrumentation, sensor | Drift, stiction, signal loss, wrong range, failed actuator or loss of communications | Transmitters, control valves, actuators, controllers, network segments, field devices | As-found calibration error, stuck travel, dead channels, communication fault logs | Yes, but only by deliberate testing. Many of these modes are hidden |
| Software and configuration | Logic error, wrong setpoint or parameter, failed update, expired certificate, corrupted database | BMS and SCADA, PLC and VFD parameters, protection relay settings, integration interfaces | Change records, event and audit logs, configuration backups diverging from the as-built | Partly. Configuration baselines and change control, not physical inspection |
Mechanisms overlap in practice; a single failure commonly involves two or three acting together.
Two observations about that table are worth drawing out. First, the widely repeated industry assertion that most rolling-element bearing failures trace back to lubrication is exactly that, an assertion that circulates without a source anyone can check. I would not put a figure on it, and neither should a report. What is defensible is the practical point behind it: lubrication-related modes are common, they are cheap to detect, and they are frequently mismanaged, so they deserve attention regardless of what proportion of failures they represent.
Second, the detectability column is doing more work than it appears. It is the column that decides whether a proactive task is even possible, and it is the honest place to record that for some modes the answer is no.
7. The mode category everyone leaves out
Software and configuration failure deserves its own paragraph, because it is commonly absent from failure-mode lists in maintenance literature, and that absence is increasingly indefensible. Modern plant fails in ways that have nothing to do with metal. A chiller sequence that never hands over to the lead unit because a schedule was edited during commissioning and never reverted. A protection relay left on default settings after a firmware update. A variable-speed drive whose acceleration ramp was changed during a fault investigation and never changed back, quietly stressing a coupling for two years. An integration interface that stops delivering meter readings because a certificate expired, so the condition-based tasks it feeds silently stop being generated.
Every one of those is a genuine failure mode with a real functional consequence, and most of them are hidden failures as well: nothing announces them, and they are found either by a deliberate check or by the eventual second event. They also break the habits of a maintenance department, because the evidence is in change logs and configuration backups rather than on a fracture surface, and the failure-finding task looks like a configuration audit rather than a test.
My recommendation is straightforward. When you build a failure-mode list for any asset with a controller, include at least the configuration and setpoint modes, the communications loss mode, and the "logic present but wrong" mode. Then give them the same treatment as everything else: an interval, a defined check, and a recorded result. A configuration baseline held under change control is the maintenance task for this mechanism, and I have very rarely seen it in a PM schedule.
8. Common failure modes by equipment class
The mechanism view is the engineering view. The equipment-class view is the one a planner can use directly, because it maps onto how asset registers are actually organised. This is deliberately a shorter treatment, and it is a starting point for a workshop rather than a finished list for any specific site.
| Equipment class | Commonly recurring failure modes | Dominant mechanisms and notes |
|---|---|---|
| Rotating equipment pumps, fans, motors, compressors, gearboxes |
Bearing spalling; bearing seizure from lubricant loss; shaft fatigue cracking; coupling wear and loosening; impeller wear opening clearances; rotor imbalance; mechanical seal face leakage; gear tooth pitting; belt elongation and fracture | Wear, fatigue, lubrication, misalignment. The richest class for condition monitoring because most modes develop progressively and show in vibration or oil |
| Static equipment and pipework vessels, exchangers, tanks, piping, valves |
Wall thinning to perforation; pitting through-wall; flange gasket leakage; valve seat passing; valve stem sticking; tube bundle fouling; tube leak; support failure allowing stress; external corrosion under insulation | Corrosion, erosion, fouling, sealing. Long development times, so inspection intervals are usually measured in months or years |
| Electrical distribution switchgear, transformers, cables, panels |
Termination loosening causing local heating; insulation breakdown to earth; contact erosion increasing resistance; partial discharge and tracking; breaker fails to trip on demand; breaker fails to close; transformer cooling loss; earthing continuity loss | Electrical and thermal. Note how many are hidden: protection that fails to operate is undetectable until demanded |
| HVAC plant chillers, AHUs, pumps, cooling towers, FCUs |
Refrigerant loss through joints; condenser and coil fouling reducing capacity; filter blockage restricting airflow; control valve sticking; damper actuator failure or linkage detachment; compressor fails to start; condensate drain blockage; belt slip; fill fouling and water treatment excursions | Fouling, sealing, controls. Degraded modes dominate here, so precise performance standards matter more than usual |
| Controls and instrumentation transmitters, controllers, BMS, networks |
Calibration drift; sensor fails to a plausible but wrong value; impulse line blockage; loss of communications on a segment; controller output stuck; wrong setpoint or parameter after a change; configuration diverges from as-built; power supply or battery backup failure | Instrument drift, contamination, software and configuration. The highest concentration of hidden modes of any class |
Indicative only. Actual modes depend on duty, environment, design and operating regime.
If you need a more rigorous starting point than a table like this, ISO 14224:2016, "Petroleum, petrochemical and natural gas industries: Collection and exchange of reliability and maintenance data for equipment", is the document to buy. It is directly relevant to this article because it provides equipment-class taxonomies together with failure-mode definitions, and it standardises how reliability and maintenance data is recorded so that it can be compared across sites and organisations. It grew out of the OREDA work, and it is worth being precise about the distinction: OREDA is a proprietary members-only database, not a standard, and citing it as one is a common error. ISO 14224 is the standard; the database is something else. It is also written for the oil and gas sector, so a facilities or utilities reader adapts rather than adopts it wholesale.
9. What failure modes are actually used for
A failure-mode list is not an end product. It is an input to four or five downstream decisions, and knowing which ones tells you how much rigour the list needs.
- Criticality and maintenance strategy. Consequence is a property of the mode, not of the asset, which is why criticality assessment done purely at asset level is coarse. The same pump can carry one mode with a safety consequence and another with only an economic one. The equipment criticality analysis guide covers the ranking; the point here is that the mode list is what feeds it.
- Deciding whether a proactive task is even possible. For each mode, ask whether there is a detectable warning, and whether the warning arrives far enough ahead to act. If the answer is no, the legitimate conclusions are redesign, added redundancy, a failure-finding task where the mode is hidden, or an accepted run-to-failure. Concluding that no proactive task is worthwhile is a real finding, not a gap in the analysis, and you should be willing to write it down.
- Designing failure code sets. The engineering concepts in this article are what a code set encodes, but the design of the code set itself, the three-level problem, cause and action structure, the trade-off between technician compliance and analytical depth, is a separate discipline covered in the failure codes guide. Build the concepts first, then let the code set follow them, because a code list designed without an underlying mode list ends up as a shopping list of symptoms.
- Feeding FMEA and RCM. The mode list is the raw material for structured analysis. The method, severity, occurrence and detection scoring, action priority and the worksheet mechanics, belongs to the FMEA guide, and the quantitative side, fitting distributions and modelling remaining useful life, to RUL and FMEA modelling.
- Structuring investigations. When something has failed, the mode is the anchor that stops an investigation drifting. Establish the mode first, then work down to mechanism and cause. The root cause analysis guide covers the techniques for getting from mode to cause.
A mode list also improves the things that sit above it. A reliability programme that cannot name the modes it is trying to eliminate is running on sentiment, which is part of why reliability engineering treats mode identification as foundational rather than optional, and why any serious attempt at improving equipment reliability starts by finding out how the equipment actually fails rather than how often.
10. Where the standards help, and what they are not
Several published documents carry the vocabulary and the analytical frameworks. All of them are voluntary, most are paywalled, and none of them is law by itself.
- IEC 60812:2018 (Edition 3), "Failure modes and effects analysis (FMEA and FMECA)". This is the international standard for the analysis technique. Worth noting that the title changed at Edition 3; earlier editions sat under a broader system reliability analysis title, and citing the old one is a frequent misquote.
- ISO 14224:2016 (third edition), as described above, for equipment-class taxonomies and reliability data collection.
- IEC 60050-192:2015, "International electrotechnical vocabulary: Part 192: Dependability". This is the formal home of dependability terminology. It supersedes IEC 60050-191:1990, so a reference to 60050-191 is out of date.
- EN 13306:2017, "Maintenance: Maintenance terminology". A European (CEN) standard with no ISO twin, and the document that defines the maintenance type vocabulary, preventive splitting into predetermined and condition-based with predictive as a form of condition-based, plus corrective in immediate and deferred forms.
- SAE JA1011_202411 and SAE JA1012_201108 for RCM criteria and the accompanying guide, relevant here for the hidden-failure and failure-finding logic.
- IEC 62740:2015, "Root cause analysis (RCA)". A real international standard for RCA does exist, which is worth stating because the opposite is often repeated. It describes principles and process steps and describes named techniques including the Why method and the Ishikawa diagram, and it is limited to after-the-event analysis, explicitly excluding the assignment of blame.
- IEC 31010:2019, "Risk management: Risk assessment techniques", if you need a catalogue of techniques to choose between. The designation is IEC 31010, not ISO 31010; only the withdrawn 2009 edition carried a dual prefix.
One document frequently miscited in this area is the AIAG & VDA FMEA Handbook, 1st edition, June 2019. It is an industry handbook rather than a standard, and its authority flows through automotive customer contracts rather than through any standards body. Calling it "the FMEA standard" is wrong. The standards bodies themselves are the right place to confirm current editions and titles: IEC , ISO and SAE International .
11. Where failure-mode lists break down
This is the section to read before committing a team to a large analysis exercise, because the failure patterns are consistent and all three are avoidable.
Lists grow unusably long without a criticality filter. There is no natural stopping point when enumerating modes. A single air handling unit analysed to component level will produce well over a hundred credible modes, and a mid-sized estate has hundreds of air handling units. Teams that start without a criticality filter run out of energy somewhere in the second asset class and abandon the exercise, leaving a partial list that is worse than none because it looks authoritative. Filter first, analyse the critical assets properly, and apply generic strategy to the rest without pretending otherwise.
Generic libraries do not transfer cleanly between sites. A mode library from another plant, a vendor, or a published taxonomy is a useful prompt and a poor substitute for local knowledge. The same pump model in a coastal plant room with poor ventilation, in an inland facility with clean dry air, and on intermittent duty with frequent starts will present genuinely different dominant modes. Duty, environment, water chemistry, operating regime and the quality of past interventions all reshape the list. Import the library as a checklist to argue against, not as an answer.
A mode list is only as good as the history behind it. This is the one that hurts. In most organisations I have worked with, the maintenance history is not a reliable record of how equipment failed. Work orders closed with "attended and rectified", failure codes selected because they were first in the dropdown, downtime unrecorded, cause fields blank. A mode list assembled from that history will reflect what people typed, not what broke. The honest response is to build the first list from engineering judgement and vendor documentation, mark clearly which entries are inferred rather than observed, and then improve the list as coding discipline improves. Do not present an inferred list as an evidence-based one.
What this costs, honestly
Doing this properly is workshop time with people who are needed elsewhere: an engineer, a planner, an experienced technician and an operator, for each asset class. There is no software that produces a good failure-mode list on its own, and any tool that claims to is offering you a generic library with your asset names substituted in. The tool can hold the list and link it to work management. It cannot know how your plant fails.
The idea to walk away with
Failure-mode analysis is a vocabulary discipline before it is an engineering one. Keep the four levels separate, function, mode, mechanism, cause. Write each mode so it names an item and a manner and points at a checkable mechanism. Accept that one item has several modes and that two items can produce the same functional failure. Ask of every single mode whether the crew would know it had happened, and give a failure-finding task to every one where the answer is no. Define functions with performance standards so degraded failures are visible. Do those five things and the taxonomy in the middle of this article becomes genuinely useful rather than decorative.
The reason this matters more than any method is that FMEA, RCM, criticality ranking, code-set design and reliability modelling all take the mode list as input. Every one of them amplifies whatever is in that list, including the errors. Getting the definitions right at the start is the cheapest quality improvement available anywhere in a reliability programme.
Final thoughts
If I had to pick one thing for a team to change tomorrow, it would be to stop writing functional failures in the failure-mode column. "Pump failed", "AHU not cooling", "motor tripped" are the three most common entries in the registers I review, and none of them supports a decision. Replacing them with entries that name the item and the manner takes no new software, no consultant and no capital, and it makes every subsequent analysis better.
The second thing would be the hidden-failure sweep. Walk the protective devices, the standby equipment, the alarms and the on-demand features, and ask for each one when it was last actually tested with a recorded as-found result. The gaps that exercise exposes are usually the largest unmanaged risks in the building, and they exist not because anyone decided to accept them but because a failure that never announces itself never generates a work order. That is the practical payoff of taking the vocabulary seriously.
Disclosure
Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.
Building a failure-mode library you can actually use?
Independent advisory on failure-mode definition, criticality filtering, failure-finding task design and turning a mode list into a workable code set and PM strategy. 22+ years across utilities, oil and gas, manufacturing, government and facility operations.
Book a conversationRelated reading: FMEA: failure mode and effects analysis, Failure analysis methods, Root cause analysis methods, Failure codes: Problem, Cause, Action, RCM introduction, The P-F curve explained, The bathtub curve explained, Equipment criticality analysis.
Muhammad Abbas
CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.
Work with me