Most of the FMEA worksheets I have been handed over twenty-two years of CMMS, EAM and asset management work were not analyses. They were documents. Somebody had produced a wide spreadsheet with the right column headings, filled it in alone from an equipment manual, computed a number in the last column, sorted by it, and filed the result where it was never opened again. The format was correct and the exercise was worthless. FMEA is not a template you complete; it is a disciplined conversation between people who know how the equipment behaves, and the worksheet is only the transcript.
The message up front: FMEA earns its cost in two places only. When it finds a failure mode with serious consequences that nothing currently detects, and when each finding leaves the room attached to a named owner and a change to a task, an interval, a spare or a design. Everything else in the method, including the scoring, is machinery to reach those two outcomes. If your FMEA does not change what a technician does next month, it did not happen.
1. What FMEA is, and what it is actually for
The FMEA meaning sits in the acronym, and it is worth reading slowly. Failure mode: the specific way something stops doing what it is supposed to do. Effects: what follows, for the equipment, the process, the people around it and the business. Analysis: systematic enumeration rather than a guess. FMEA is therefore an inductive, bottom-up method. You start at component or step level, name every plausible way it can fail, and reason forwards to the consequences.
That direction of travel tells you what FMEA is good at. Reasoning forwards from many causes to their effects makes it excellent at breadth: it surfaces the unremarkable failure modes nobody has thought about, which is exactly where unpleasant surprises live. It is correspondingly poor at depth and at combinations, because it looks at one failure mode at a time.
The purposes it genuinely serves, ranked for a maintenance or facilities reader:
- Finding the undetected consequential modes. The output that justifies the exercise is a short list of rows where the consequence is serious and the current detection is honestly "nothing". That list is often a surprise and usually actionable.
- Deciding what maintenance to do, how often, and what to stop doing. A task not tied to a named failure mode is a task nobody can defend or delete. And since you cannot act on forty findings, ranking them defensibly is part of the method.
- Capturing knowledge before it walks out. The technician who has repaired these pumps for fifteen years holds a failure mode catalogue that exists nowhere else. Written down, it becomes the coding structure in a CMMS, the basis of condition monitoring selection, and the input to any quantitative reliability work.
The method originated in aerospace and defence, moved into automotive quality work, and reached maintenance strategy after that. That lineage matters, because most FMEA training material and almost all the scoring conventions in circulation were written for manufacturing quality rather than for maintaining installed plant. Recognising the mismatch is the subject of section 3.
2. Where the method is actually defined
Ask five people for "the FMEA standard" and you will get five different answers, at least three of them wrong. Getting this right affects which conventions you are obliged to follow and which you are merely inheriting from somebody else's industry.
The international standard is IEC 60812:2018, Edition 3, "Failure modes and effects analysis (FMEA and FMECA)." Note the title, because it changed at Edition 3: earlier editions were titled around "Analysis techniques for system reliability", and citing that older title as current is one of the most common misquotes in reliability writing. Edition 3 covers both FMEA and FMECA and is the neutral, industry-independent description of the method. Published by the International Electrotechnical Commission , paywalled, voluntary, and not law anywhere by itself.
The AIAG & VDA FMEA Handbook, 1st edition, June 2019, is an industry handbook and not a standard. Produced jointly by the Automotive Industry Action Group and the German automotive association VDA to harmonise two previously divergent approaches, it is not an ISO, IEC or ANSI document and no standards body has adopted it. Its authority flows entirely through automotive customer contracts: if your customer requires it, you follow it because the contract says so. It is nonetheless the most influential FMEA document in circulation, and the origin of two things you will meet constantly, the seven-step process structure and the Action Priority tables that displaced the Risk Priority Number in automotive practice. Available through the AIAG .
SAE J1739_202101 is an SAE Recommended Practice, "Potential Failure Mode and Effects Analysis", deliberately aligned with the AIAG and VDA seven-step approach. This is what most commentary gets wrong: J1739 and the AIAG-VDA handbook are alternative, largely compatible expressions of the same method, not a sequence in which one replaced the other. Neither is withdrawn. Available from SAE International .
One regulatory point, and the jurisdiction matters. In the United States, FMEA is one of the methodologies named as acceptable for a process hazard analysis under the Occupational Safety and Health Administration's Process Safety Management rule at 29 CFR 1910.119(e). That is a federal United States requirement with no legal force in the United Kingdom, the European Union or the Gulf, and a process hazard analysis is a process and scenario level assessment of a covered chemical process, not a task or component level maintenance analysis. Running an FMEA on a chiller to design a preventive maintenance programme is not a PSM process hazard analysis.
The standards question matters less than people hope
None of these documents mandates a worksheet layout, and the scoring scales in circulation are conventions rather than requirements outside a contractual context. IEC 60812:2018 gives you a rigorous vocabulary and a defensible process description. It will not tell you whether your severity 7 should have been an 8. That judgement stays with the people in the room, which is why who is in the room is the real quality control.
3. Design FMEA, Process FMEA, FMECA, and which one you want
Nearly all FMEA confusion traces back to people applying a variant built for a different question.
Design FMEA (DFMEA) analyses a product or system design. The item is a component or subsystem, the failure modes are ways the design fails to deliver its intended function, and the actions are design changes. It happens before anything is built. Process FMEA (PFMEA) analyses a manufacturing or assembly process. The item is a process step, the effects run downstream to the product and the customer, and the actions are changes to process controls, fixtures, inspections and work instructions. PFMEA is the variant most FMEA training teaches, because it is the variant automotive suppliers are contractually required to produce.
Now the important part. If you are a maintenance planner or a facilities manager, neither of those is what you want. You are keeping installed equipment functioning, and your actions are maintenance tasks, intervals, spares, condition monitoring and occasionally a modification. When people force plant maintenance into a PFMEA template, they produce rows where "process step" has been quietly redefined as "asset" while the scoring scales still talk about customer dissatisfaction and defect rates. What you want is closer to one of two things:
- FMECA, Failure Mode, Effects and Criticality Analysis. The extra C adds an explicit criticality assessment, the part a maintenance reader cares about most: not just what can fail but which failures matter enough to spend money on. IEC 60812:2018 covers FMECA alongside FMEA, and for installed plant it is generally the better fit.
- An RCM style analysis, in which an FMEA step sits inside a wider decision process that also classifies failure consequences and selects task types through defined logic. This is what most maintenance organisations are reaching for when they say they want an FMEA.
Do not agonise over the label, decide what the output has to be. If it is maintenance tasks with intervals, build consequence classification and task selection in from the start and call it FMECA or RCM as suits your documentation. If it is a risk register informing an engineering decision, a straight FMEA is enough. For failure mode taxonomies themselves, see the guide to common failure mode types.
4. Define the item and its boundary first, the step almost everyone skips
This is the step that gets left out, and leaving it out is why so many workshops drift for three hours and produce forty rows nobody can use. Before you name a single failure mode, write down three things: what the item is, where its boundary runs, and what operating context you are analysing it in.
What the item is. Pinned to the asset register, using the identifier that exists in your CMMS or EAM, not a description invented in the workshop. If the FMEA says "chilled water pump 2" and the register says "CHW-P-002", the findings will not attach cleanly to job plans. Decide which level of the hierarchy you are analysing at, and stay there.
Where the boundary runs. This is the one that causes real damage when it is vague. Does the pump FMEA include the motor? The variable speed drive? The suction strainer? The isolation valves? There is no universally right answer, but there is a universally wrong one, which is not deciding. Two things follow from an undefined boundary. Items fall through the gap, because the pump team assumed the drive was electrical scope and the electrical team assumed the pump FMEA covered it. And the same modes get analysed twice with different conclusions, producing contradictory recommendations for one failure. Draw the boundary explicitly, list the interfaces crossing it, and state which side owns each.
What operating context applies. The same physical pump has different failure modes and radically different consequences depending on whether it is duty or standby, continuous or intermittent, redundant or not, and on ambient conditions. A cooling pump in an Abu Dhabi summer is a different reliability problem from the identical model in a temperate climate. Where the same equipment runs in two contexts, that is two analyses.
The half hour that saves the workshop
Spend the first thirty minutes agreeing the item, the boundary and the operating context, written on the wall where everyone can see it, and do not name a failure mode until that is done. Workshops that run badly are commonly running without it. Half an hour, and it prevents the two most expensive defects in the output.
Which items to analyse at all is a prior question, answered by a criticality screen. You cannot FMEA an entire estate, and attempting it produces a shallow analysis of everything instead of a useful analysis of what matters. Rank first, then analyse the top of the list: the equipment criticality analysis guide covers that screen.
5. The anatomy of a row, done properly
An FMEA row is a chain of reasoning, and it is only as strong as its weakest link. The links, in order, are function, failure mode, effect, cause, current controls and detection. One of the most consequential errors in the whole method is a vague failure mode, because the mode is the hinge on which everything after it turns.
Consider what happens when somebody writes "pump fails". What is the effect? Unanswerable: a seized pump and a pump whose output has fallen twenty percent produce entirely different effects. What is the cause? Unanswerable, because there are thirty of them. What detects it? Unanswerable, because vibration analysis finds a spalling bearing and does not find a blocked strainer. All three scores become guesses and the action, if one survives, is a generic "increase inspection". A properly stated mode, for instance "impeller wear reducing delivered head", makes every subsequent column answerable.
The table below is the row structure I would use: the column, what it must contain, and what a bad entry looks like. The third column is deliberate, because most people learn this method faster by recognising their own habits than by reading a definition.
| Column | What it must say | What a bad entry looks like |
|---|---|---|
| Item and boundary | The asset or component identifier as it exists in the register, at a stated hierarchy level, with the boundary already agreed | A description invented in the workshop that matches nothing in the CMMS |
| Function and performance standard | A verb, an object and a measurable standard in the stated operating context, so that failure is a departure from something written down | Naming the equipment instead of the duty, for instance "cooling pump" with no flow, head or availability standard |
| Failure mode | One specific physical mechanism or state at component level, granular enough that a task could be designed for it and only it | "Pump fails", "does not work", "breakdown". A category, not a mode |
| Failure effect | What follows, concretely: what an operator would observe, secondary damage, safety and environmental exposure, operational and production impact, repair demanded | "Pump stops." A restatement of the mode rather than its consequences |
| Failure cause | Why the mode occurs here, at a level where a countermeasure exists. The same mode with two causes is two rows if the countermeasures differ | Either repeating the mode, or drilling five layers down into organisational root cause, which is a different method |
| Current controls | What exists today that prevents the cause or reveals the developing mode: specific PM tasks, condition monitoring routes, alarms, operator rounds, trips, or honestly "none" | Crediting controls that exist on paper but are not actually performed, or listing a control that cannot detect this mode |
| Detection | Whether existing controls would reveal the mode with enough warning to act, judged against how fast this mode develops | Assuming any monitoring equals good detection, regardless of interval or of whether the mode gives any warning at all |
| Rating and priority | Severity, occurrence and detection ratings on defined scales, plus a priority reached by a stated rule | Numbers with no scale definition attached, and a threshold applied to a product |
| Action, owner, date | A specific change to a task, interval, spare, control, instrument or design, with one named person and a date | "Monitor closely." No owner, no date, no verb a planner can act on |
Two habits are worth building. Keep cause and mode strictly separate and resist chasing causes to their ultimate root: FMEA is a breadth instrument, and when you need depth on a single failure you switch tools, which is what root cause analysis methods are for. Second, be honest in the current controls column, because this is where the analysis either earns its cost or becomes theatre. A monthly vibration route actually executed quarterly does not deserve credit for monthly detection. The rows where the honest answer is "nothing detects this" are the rows the exercise exists to find, and flattering this column is how organisations bury them.
6. A short worked extract
Below is a small illustrative extract, using a generic equipment example of my own construction rather than any real installation. Treat the ratings as teaching material, not as values to copy.
Item: illustrative chilled water distribution pump, hypothetical designation CHW-P-001. Boundary: pump casing, impeller, shaft, bearings, mechanical seal and coupling. Motor, variable speed drive, suction strainer and isolation valves explicitly excluded and analysed separately. Operating context: duty pump in a duty and standby pair, continuous operation, hot ambient. Function: deliver chilled water to the distribution header at design flow and differential pressure whenever the plant is called.
| Failure mode | Effect | Cause | Current controls and detection | Illustrative S / O / D |
|---|---|---|---|---|
| Bearing spalling, drive end | Rising vibration and noise, then shaft and seal damage if run on; standby pump carries load; unplanned intervention and extended repair | Inadequate or contaminated lubrication; misalignment induced load | Monthly vibration route; develops over weeks, so detection is genuinely good if the route is performed | 7 / 4 / 3 |
| Mechanical seal face wear leading to leakage | Progressive leak to plant room floor, water damage and slip hazard, bearing contamination if unaddressed | Dry running during a low flow condition; abrasive particles in the fluid | Operator round records visible leakage; detected only once leakage is visible, giving limited warning | 5 / 6 / 5 |
| Impeller wear reducing delivered head | No alarm and no visible symptom; distribution temperatures drift, comfort complaints attributed elsewhere, energy consumption rises as the system compensates | Erosion from entrained particulate over years of service | None. No differential pressure or flow trend is recorded, so the degradation is invisible | 6 / 5 / 9 |
| Coupling element failure | Immediate total loss of pump output; standby starts; minimal secondary damage; short repair | Elastomer ageing in hot ambient; residual misalignment | Annual inspection at overhaul; failure is effectively sudden, so there is little warning to detect | 4 / 3 / 8 |
Read the third row again, because it is the reason to run the analysis. A slowly degrading impeller with no instrumentation never appears in a failure report and never triggers an alarm, and it quietly costs energy and comfort for years while being blamed on the chillers or the control strategy. The action it generates is small and concrete: record and trend pump differential pressure, set a review threshold. That is an FMEA doing its job, and it came from the detection column being answered honestly rather than from the arithmetic.
The fourth row makes the opposite point: a sudden mode with a standby behind it and a short repair is a candidate for accepting the failure and holding the spare. Whether a mode develops gradually or suddenly determines whether condition based detection is possible at all, and the P-F curve explained makes that precise. A full walkthrough with ratings justified line by line and the arithmetic set out belongs in the companion piece, an FMEA example step by step with RPN calculation.
7. Who must be in the room
An FMEA done alone is worthless, and I mean that literally. The method's value comes from combining knowledge no single person holds: what the equipment is supposed to do, how it behaves on this site, what has really failed, what is really being maintained, and what losing it costs the business. One person working from drawings and a manual produces a document containing none of that, and its most dangerous property is that it looks like an analysis. The composition I would insist on:
- A facilitator who owns the method, not the content. Keeping the boundary honest, stopping the group solving problems mid analysis, refusing vague failure modes. Not the person with the strongest opinions about the equipment.
- Someone who operates it on shift, not a supervisor describing how it ought to run, but the person who has watched it misbehave at three in the morning.
- A technician who repairs it. The most valuable participant and the one most often cut on cost grounds.
- A reliability or engineering voice, arriving with the maintenance history already extracted and analysed rather than recalled.
- A planner, who owns the current controls column honestly and who will convert every recommendation into a job plan in whatever system you run, IBM Maximo, SAP PM, Hexagon EAM, Planon or a mid market CMMS.
- Specialists on call for the sections needing them: controls, electrical, safety, process. An hour each, not the whole study.
Six to eight people is the working maximum, and two to three hour sessions beat full days, because the rows produced in hour six are the ones you will regret. Bring the maintenance history as a document on the table rather than as recollection, which over weights the dramatic failure everybody remembers and forgets the frequent minor one that costs more.
The readability test
Hand a finished row to a competent engineer who was not in the room. If they independently reach the same consequence judgement and the same recommended action, it is written well enough to survive. If they have to ask what you meant, it will be unusable in three years. Apply that test to a sample of rows before closing the study; it catches more real defects than score review.
8. Severity, occurrence and detection: scoring honestly
The scored form of FMEA rates each row on three scales, conventionally one to ten, and each has a characteristic way of being got wrong.
Severity rates the consequence of the failure effect. It is a property of the effect alone, and the most common error is contaminating it with frequency. "It is severe but it hardly ever happens, so call it a 5" is a corrupted score that double counts rarity, because frequency is what occurrence is for. If your organisation has a corporate risk matrix with defined consequence categories for safety, environment, production and cost, make your severity scale be that matrix: one language for consequence across the business, and no workshop inventing its own.
Occurrence rates how frequently the cause produces the mode in this context, and it is the score that can be evidenced rather than debated. If you have recorded failures of this mode across a population of similar assets over several years, occurrence is a matter of data, and bringing that data into the room changes the discussion entirely. Where it does not exist, say so on the row rather than quietly guessing: a flagged estimate is useful, an unflagged one is misleading. This requires failure history coded consistently enough to count, a prerequisite most organisations discover they do not meet.
Detection rates the likelihood that existing controls reveal the developing failure in time to act, and it is the most consistently over scored of the three. Teams credit controls that exist on paper rather than controls performed. And teams ignore the interval question: a technique whose interval exceeds the time between the failure becoming detectable and the failure occurring cannot detect it in time, however good the technique. A monthly route will not catch a mode that develops in ten days. Scoring detection without asking how fast the mode develops is the biggest source of false comfort in FMEA output.
9. The Risk Priority Number, and what has replaced it
Multiply severity by occurrence by detection and you get the Risk Priority Number, a value from 1 to 1000. Teams set a threshold, act on everything above it, close everything below. It is popular because it is simple, sortable and easy to report upward. As a risk measure it is also indefensible, and there are three separate problems rather than one.
It multiplies ordinal scales. The ratings are ordinal: a severity of 8 is worse than a 4, but nothing establishes that it is twice as bad, and nothing guarantees the gap between 6 and 7 equals the gap between 2 and 3. Multiplication is defined on ratio scales, where numbers carry real magnitudes. Multiplying ordinal ranks produces a value with no meaning: you can compute it, you cannot interpret it.
Equal Risk Priority Numbers can mean entirely different risks. An RPN of 120 might be 10 times 4 times 3, a rare but potentially fatal failure with reasonable detection. Or 2 times 6 times 10, a trivial consequence that happens often and is never detected. Those demand completely different responses, but the number cannot tell them apart and a threshold treats them identically.
A high severity item can hide below the threshold. This is the failure that hurts people. A mode with severity 10, occurrence 2 and detection 2 gives a product of 40, below any conventional threshold, so the row closes with no action. A potentially fatal failure mode has been filtered out by arithmetic, while three moderate, frequent, poorly detected modes score above the line and consume the action budget. Sorting by RPN systematically de prioritises exactly the rows a safety conscious organisation should look at first.
What to do instead, and what it costs you
Two changes, in order of value per unit of effort. First, a severity gate applied before any arithmetic: any mode above a defined severity gets an action regardless of its other two scores. One rule, no cost, and it fixes the most dangerous defect. Second, replace multiplication with a lookup: a defined high, medium or low priority for each meaningful combination of severity, occurrence and detection, evaluated in that order of importance. This is the Action Priority approach introduced in the AIAG & VDA FMEA Handbook and adopted in SAE J1739_202101, now standard practice in automotive FMEA. The honest cost is convenience: you lose the single sortable number that made RPN easy to put in a steering committee pack, and your team must defend a priority judgement rather than point at a threshold. If your organisation will not accept that trade, at least implement the severity gate.
One further caution: FMEA scoring ranks failure modes within a system already selected for analysis. It says nothing about how important that asset is relative to others, and it should never be used as an asset criticality ranking.
10. Turning findings into maintenance tasks, which is where most FMEAs die
This is the section the method's literature covers least well and the one that decides whether the investment returns anything. An FMEA that ends with a prioritised list of findings has not finished; the finished product is a change to what people do. For each finding you intend to act on, decide which type of response fits the mode. The options are genuinely limited, which makes the decision faster:
- A condition based task, where the mode gives detectable warning and you can monitor at an interval comfortably shorter than the warning period. The preferred answer when available, because it intervenes on evidence rather than the calendar.
- A fixed interval restoration or replacement, where the mode has a reasonably predictable wear out pattern and no usable warning. The interval needs evidence, which is where quantitative reliability modelling earns its place.
- A failure finding task, where the mode is hidden: a standby pump that will not start, a protective device that will not trip, a fire damper that will not close. These produce no symptom until they are needed, so the only possible task is periodic function testing. Hidden failures are routinely missed by inexperienced teams and are among the most valuable findings.
- A design or operating change, where no task is effective: adding instrumentation or redundancy, changing a material, altering the operating regime.
- An accepted failure with no scheduled task, where the consequence is tolerable and managing it costs more than suffering it. This must be a written decision with the reasoning recorded. Accepted by decision is legitimate; accepted by neglect is not the same thing.
Then the mechanical part, where the failure usually occurs. Each accepted action needs a named individual, a date, and a specific destination in a system. "Add a quarterly differential pressure reading to job plan CHW-PM-Q and record the value against the asset" is implementable. "Improve pump monitoring" is not. Every action should end up as one of a small set of concrete artefacts: an amended job plan task line, a changed interval, a new condition monitoring route point, a stocked spare, an instrument on a capital list, or a documented acceptance. For how these outputs land in a programme, see the complete guide to preventive maintenance.
Two things kill this stage. Findings written in a workshop vocabulary nobody can map onto a job plan, which is why the planner needs to be in the room. And the absence of a review cycle, so the worksheet becomes historical. A finished FMEA needs a scheduled revisit triggered by time and by events: a significant failure that was not on the worksheet, a modification, a change of duty, a change of maintenance provider.
Where findings must become intervals with evidence behind them rather than intervals somebody chose, the analysis has to go quantitative: fitting distributions to observed failure data, handling the censored records that make up most maintenance history, and producing remaining useful life estimates with honest bounds. That is a separate body of technique, covered in from RUL to FMEA: modelling equipment failure, which pairs the qualitative analysis here with the statistical modelling that turns a prioritised mode into a defensible interval. This guide stays deliberately on the method side of that boundary.
11. FMEA inside RCM
This relationship is simpler than the literature makes it look. Reliability Centred Maintenance is a wider decision process with an FMEA style step inside it. RCM works through functions, then functional failures, then failure modes and effects, which is the FMEA content, and adds two things FMEA does not require: explicit classification of failure consequences including the hidden failure category, and a defined task selection logic.
The governing document is SAE JA1011_202411, "Evaluation Criteria for Reliability-Centered Maintenance (RCM) Processes", published November 2024. Be precise about what it does: it does not prescribe a process. It sets the criteria defining what may legitimately be called RCM, as questions every RCM analysis must answer in order, plus required consequence and task selection logic. A process failing any criterion is not RCM whatever the marketing says, which is a useful test for anything sold under the name. SAE JA1012_201108 is a guide to it, now one generation behind. The practical implication: if your intended output is maintenance tasks, running the analysis within an RCM framework gets you the consequence classification and the task selection logic that a bare FMEA leaves you to invent. See the introduction to RCM.
12. When FMEA is the wrong tool
The house rule in my writing is to say where a method does not work, and FMEA has clear limits.
- When the failure involves combinations. FMEA in its standard form considers one mode at a time. If your concern is a failure requiring two or three things to fail together, for instance a protective system that only leaves you exposed when the primary control has also failed, you need a deductive top down method working backwards from the undesired outcome through logic gates. That is fault tree analysis, and the complete guide to fault tree analysis covers where it fits alongside FMEA.
- When something has already failed and you need to know why. FMEA is prospective and broad; investigating a specific event demands depth, evidence handling and causal reasoning rather than enumeration.
- When the dominant mechanism is human or organisational. Failures rooted in decision making, competence, supervision or commercial pressure fit awkwardly into a row, and forcing them in produces no meaningful mode and no implementable action.
- When you cannot get the people. If the operator and the technician cannot be released, do not run the workshop and produce a document from manuals instead. A gap in coverage is honest. A document that looks like an analysis and is not is worse than nothing, because it will be trusted.
- When the asset does not justify it. A thorough FMEA on a cheap, low consequence, easily replaced asset misallocates your scarcest resource, the attention of the people who know the equipment.
And one general limitation: FMEA finds only the failure modes the people in the room can imagine, which is an argument for revisiting the worksheet every time reality produces a mode you did not predict.
The idea to walk away with
FMEA is a conversation with a filing system attached, and almost every way it fails comes from treating the filing system as the point. The value concentrates in three places: agreeing the item, its boundary and its operating context before anything else; stating failure modes specifically enough that effect, cause and detection all become answerable; and being honest in the current controls column so the rows where nothing detects a serious consequence actually surface. Get those three right and the analysis produces findings worth acting on even if the scoring is rough. Get the scoring perfect and those three wrong, and you have an elegantly quantified document describing equipment that does not exist.
Final thoughts
If you are being asked to produce an FMEA and want it to be worth the time, change what you optimise for. Do not aim for a complete worksheet. Aim for a small number of findings that change what happens next month, each with a named owner, a date and a specific destination in your maintenance system. Ten rows that alter job plans are worth more than four hundred that were filed.
Start with one critical asset, spend the first half hour on item, boundary and context, insist on the operator and the technician being in the room, be brutal about vague failure modes, and be honest where the detection answer is nothing. Then convert each accepted finding into an amended task, a changed interval, a stocked spare or a written acceptance. That is the whole method. Everything else is formatting.
Disclosure
Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.
Setting up an FMEA or FMECA programme?
Independent advisory on failure mode analysis, criticality screening, task selection and turning findings into job plans that actually land in Maximo, SAP PM, Hexagon EAM or a mid market CMMS. 22+ years across utilities, oil and gas, manufacturing, government and facility operations.
Book a conversationRelated reading: FMEA example step by step with RPN calculation, From RUL to FMEA: modelling equipment failure, Failure modes: common types and how to analyse them, Fault tree analysis: a complete guide, Root cause analysis methods, Equipment criticality analysis, The P-F curve explained, Introduction to RCM, Preventive maintenance: the complete guide.
Muhammad Abbas
CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.
Work with me