Most people looking for an FMEA example have already read the theory and are stuck on the blank worksheet. They know the columns. What they do not know is what a defensible row actually looks like: how specific a failure mode has to be before it is useful, how you defend a severity of 7 rather than 8, and what you are supposed to do with the risk priority number once you have multiplied the three ranks together. This article does not re-teach the method. It works one example all the way through, shows the judgement behind every score, and is honest about what the RPN can and cannot tell you.
The message up front: every asset, score and number on this page is invented for teaching. Do not copy a single value into your own analysis. Severity, occurrence and detection are judgements about your equipment, your operating context, your controls and your consequences, and a number lifted from someone else's worksheet is worse than no number at all because it looks authoritative. Copy the shape of the reasoning, never the figures. If you want the method, the standards landscape and the design versus process FMEA distinction, that sits in the FMEA guide. This page is the worked example.
1. The case: one asset, clearly labelled hypothetical
I am going to analyse a chilled water pump set, because pumps are the asset class most facilities and maintenance teams already understand, and because they mix gradual and sudden failure modes in a way that makes the scoring genuinely interesting rather than mechanical. The asset, for the rest of this article, is CHW-P-01, an electrically driven end-suction centrifugal pump forming the lead unit of a duty/standby pair on the primary chilled water loop of a mid-rise commercial office building. It is invented. So is the building, so is the duty, so are the performance standards, and so is every single score. There is no such thing as a standard severity score for impeller wear.
The order of work below is the order that produces defensible rows: define the item and its boundary, state the functions properly, derive the functional failures, then and only then reach for failure modes. Teams that start at failure modes produce a list of things that can break rather than an analysis of risk, and they usually cannot explain afterwards why anything on the list matters.
2. Step one: the item, the boundary and the operating context
The boundary decision is the one most often skipped, and it silently determines half the content of the worksheet. If the motor is inside the boundary, winding failure is a failure mode of your item. If it is outside, winding failure becomes an external cause and the analysis shifts. Neither is wrong. What is wrong is not deciding, because then two people in the workshop are analysing different machines. For CHW-P-01 I am drawing it as follows, and writing it at the top of the worksheet so the reviewer three years from now knows what was and was not considered.
- Inside the boundary: pump casing, impeller, wear rings, shaft, mechanical seal, bearings, coupling, baseplate and grout, motor, local starter and variable frequency drive, suction strainer, and the pump set's own instrumentation.
- Outside the boundary: the chilled water generation plant, the distribution pipework beyond the isolation valves, the building management system head end, the electrical supply upstream of the starter, and the standby pump.
- Interfaces to name explicitly: the BMS duty/standby changeover command, the incoming electrical supply, and the loop water chemistry, which is controlled by someone else but drives at least two of my failure modes.
Operating context matters as much as the boundary, because the same pump in a different context gets different scores. Mine, illustrative: continuous operation during occupied hours, an automatic standby unit that is tested but whose changeover has never been proven under load, a punishing ambient design condition for several months of the year, and a tenant population who will call the help desk within twenty minutes of space temperature drifting. That last point is why several severity scores below are higher than a purely mechanical reading would suggest.
Why context changes everything
A duty/standby pair with a proven automatic changeover has a lower severity on loss of flow than a single pump, because the functional consequence is degraded redundancy rather than lost cooling. The moment you cannot prove the changeover works, that severity reduction evaporates. Write your redundancy assumption into the worksheet, because an auditor will ask what it was based on.
3. Step two: state the functions properly
A function is not "pump water". That describes the machine rather than stating what it must achieve, and it cannot be failed in any useful way. A function has a verb, an object and a performance standard, so failure becomes a factual question rather than an opinion. Written properly, and with all figures illustrative:
- F1: Deliver chilled water to the primary loop at not less than 95 litres per second against 28 metres of head whenever the pump is called to lead duty.
- F2: Contain the chilled water within the pressure envelope with no visible leakage to the plant room floor.
- F3: Start and reach stable duty within 30 seconds of receiving a lead or changeover command from the BMS.
- F4: Operate with bearing housing vibration and casing temperature within the limits recorded in the commissioning acceptance record.
- F5: Modulate flow on the variable frequency drive across the operating band in response to the loop differential pressure setpoint.
F2, F3 and F4 are the functions people routinely forget, because they are not the headline duty, yet containment, start and protective-function failures generate a large share of real maintenance pain. F4 is a particularly useful habit: an asset running outside its acceptance envelope has already failed a function, which is exactly the framing that justifies intervention before the breakdown. That link between functional failure and condition evidence is developed further in the RUL and failure modelling article.
4. Step three: derive the functional failures
Each function fails in more than one way, and the partial failures matter as much as the total ones. A functional failure is simply a statement that a performance standard is not being met.
- FF1 (from F1): delivers flow, but below 95 litres per second at the required head.
- FF2 (from F1): delivers no flow at all.
- FF3 (from F2): leaks chilled water to the plant room.
- FF4 (from F3): does not start, or does not reach duty within 30 seconds, on command.
- FF5 (from F4): operates with vibration or temperature outside the acceptance limits.
- FF6 (from F5): does not modulate, running fixed speed or tripping on the drive.
Splitting FF1 from FF2 is not pedantry. Partial and total loss of flow have different effects, different detection prospects and different scores, and merging them into "pump failure" destroys the analysis. Gradual, detectable modes land in the partial row, which is where condition monitoring earns its place. Sudden modes land in the total row, and no monitoring interval saves you from those.
5. Step four: failure modes, specific enough to act on
A failure mode names the physical mechanism, not the symptom. "Bearing failure" is a symptom of at least four mechanisms, each with a different score, control and action. If you cannot write a plausible task against a mode, it is not specific enough yet. The broader taxonomy, and the trap of confusing mode with effect and with cause, is covered in common failure mode types and how to analyse them.
I could have written forty modes. Eleven is what a workshop of the right people can score honestly in one session, and a short analysis that gets acted on beats an exhaustive one that gets filed.
- M1: Impeller and wear ring erosion from particulate in the loop water (FF1)
- M2: Discharge isolation valve left partially closed after maintenance (FF1)
- M3: Suction strainer progressive blockage (FF1 then FF2)
- M4: Motor winding insulation breakdown (FF2)
- M5: Flexible coupling element fracture (FF2)
- M6: Mechanical seal face wear (FF3)
- M7: Casing joint gasket relaxation (FF3)
- M8: BMS changeover interposing relay failure (FF4)
- M9: Starter contactor contact welding (FF4)
- M10: Bearing degradation from lubrication starvation (FF5)
- M11: Grout and holding-down deterioration causing coupling misalignment (FF5)
6. Step five: effects at local and system level
Severity is scored against the effect, so the effect has to be written at the level where the consequence actually lands. Two levels are the minimum: local, meaning what happens to the item itself, and system, meaning what the building, the process or the occupant experiences. Worksheets that record only the local effect end up with flat, uninformative severity scores.
Take M5, coupling element fracture. Locally: drive is lost, the motor runs free, the pump stops delivering. At system level: the lead pump is offline and the loop depends entirely on an automatic changeover whose reliability is exactly what M8 and M9 are about. Then the end effect, worth naming when it applies: if the changeover also fails, cooling is lost to occupied floors at peak ambient, which becomes a tenant complaint, a service credit exposure and possibly a risk to sensitive equipment rooms. Severity is scored against the worst credible end effect, not the tidy local one.
7. Step six: current controls and detection
Detection is the most commonly abused column, because teams score the detection they intend to have rather than the detection they actually have. Write the existing control down first, by name, then score against it. If the control is "the operator would probably notice", write that.
The illustrative current controls on CHW-P-01: a monthly operator round with a visual and audible check, a six-monthly PM task covering lubrication and general inspection, BMS monitoring of loop differential pressure and pump run status, a drive fault alarm, an annual thermographic survey of the electrical panel, and a monthly automatic changeover test that is scheduled but whose result is recorded only as a tick. That last detail is doing a lot of work in the scores below, and it only surfaces if you ask to see the completed record rather than the schedule.
Where this exercise is weakest
An FMEA scored by one person at a desk, which is effectively what this article is, is a teaching artefact and not a risk assessment. The scores only become defensible when an operator, a maintenance technician who has actually opened this pump, a controls engineer and someone who owns the consequence argue them out in the same room and the disagreements get recorded. My single-author numbers below have no such validation, which is one more reason not to copy them.
8. Scoring scales are a local design choice
There is no authoritative universal scale. Scales are a local design decision, and their only real job is to make two people in the same organisation mean roughly the same thing by a 7. A three-point or five-point scale is often enough for maintenance work; ten-point scales come from manufacturing practice and give a false impression of resolution. What matters more than the number of points is that every point has a written anchor tied to your own consequences, your own frequency data and your own controls.
Below is an illustrative ten-point shape, offered purely as something to adapt. The anchors are mine, invented for this article, and no scale anywhere is authoritative. Rewrite them in your own language and get them signed off before the first workshop rather than argued about during it.
| Band | Severity anchor (illustrative) | Occurrence anchor (illustrative) | Detection anchor (illustrative) |
|---|---|---|---|
| 1 | No discernible effect on function or service | Failure essentially eliminated by design | Control detects the mode before any loss of function, reliably |
| 2 to 3 | Minor degradation, noticed by maintenance only | Isolated occurrences on similar assets | Strong chance of detection with clear warning time |
| 4 to 6 | Service degraded, occupants or process affected, no safety impact | Occasional, credible within the asset life | Moderate chance of detection, warning time uncertain |
| 7 to 8 | Loss of primary function, service interruption, regulatory or contractual exposure | Repeated occurrences expected | Weak detection, usually found after loss of function |
| 9 to 10 | Safety or statutory consequence, or loss with no workaround | Failure almost inevitable in service | No credible detection before failure |
One structural point: severity is a property of the effect, so it does not move when you improve inspection. Only a design, redundancy or consequence change moves severity. Occurrence moves when you reduce the cause. Detection moves when you improve the control. That discipline stops the common self-deception where a team adds an inspection and quietly drops the severity to get under a threshold.
9. Step seven: scoring, with the reasoning shown
This is the part that teaches, so I am going to narrate the judgement rather than just assert numbers. Every figure is illustrative.
M1, impeller and wear ring erosion. Severity 6: the effect is progressive loss of head and flow, so cooling capacity degrades and occupants notice on hot days, but there is no safety consequence and the pump keeps running. Occurrence 5: loop water chemistry is managed by another party and side-stream filtration is absent, so erosion is a credible recurring mechanism rather than a rarity. Detection 4: pump differential pressure is trended on the BMS, so a slow decline is visible to someone who looks, but nobody currently owns the trend review, which is what stops this being a 2.
M2, valve left partially closed. Severity 5: reduced flow, same character as M1 but usually caught quickly. Occurrence 4: human error after intrusive work is common wherever valve line-up is not a signed-off step. Detection 3: the shortfall appears immediately on the BMS and the timing correlates with a completed work order.
M3, suction strainer blockage. Severity 7: progressive blockage can end in cavitation and total loss of flow, which puts it in the loss-of-primary-function band. Occurrence 4: credible on a loop with known particulate. Detection 4: strainer differential pressure is measurable but not alarmed, so detection depends on the monthly round.
M4, motor winding insulation breakdown. Severity 8: immediate total loss of duty with no warning and a long repair or replacement lead time. Occurrence 2: genuinely infrequent on a correctly sized motor in a clean dry plant room. Detection 6: there is no insulation resistance or polarisation index trending in place, so in practice this is found when the motor fails, and the annual thermographic survey of the panel does not look at the windings. Hold on to this row. It becomes the most important argument in the next section.
M5, coupling element fracture. Severity 8: same end effect as M4, sudden total loss of drive. Occurrence 3: elastomeric elements do degrade, and misalignment accelerates it. Detection 5: visible on inspection if the guard is removed, which the six-monthly task does not currently require.
M6, mechanical seal face wear. Severity 4: nuisance leakage, housekeeping and corrosion issue, no loss of duty in the short term. Occurrence 6: seals are consumables and this is the single most frequently recurring mode on the asset. Detection 3: visible water on the floor, found on the monthly round.
M7, casing gasket relaxation. Severity 5, occurrence 3, detection 4. Similar to M6 but less frequent and less obvious until the leak is established. M9, contactor contact welding. Severity 7, occurrence 3, detection 6. Related to M8 in effect, slightly more findable because contact condition can be inspected and thermography sometimes catches a hot contactor. M11, grout and holding-down deterioration. Severity 5, occurrence 3, detection 5. Slow, insidious, and rarely looked for because baseplate condition is not on anybody's checklist.
M8, changeover interposing relay failure. Severity 7: on its own this does nothing at all, which is exactly the problem. It only shows up when the lead pump fails and the standby does not pick up, and then the effect is full loss of cooling. Occurrence 3: relays fail, and this one is unmonitored. Detection 8: this is a hidden failure. The scheduled changeover test exists but the result is recorded as a tick with no witnessed evidence, so there is effectively no reliable detection. Hidden failure modes are where high detection scores belong, and they are the modes most often missed entirely because nothing visible happens when they occur.
M10, bearing degradation from lubrication starvation. Severity 6: left alone it progresses to seizure, but it gives warning. Occurrence 5: lubrication practice is the most common maintenance-induced weakness I see on pump sets. Detection 4: audible and thermal symptoms appear on the round, but there is no vibration route, so detection is late relative to the available warning period.
10. The full worked FMEA worksheet
Here is the whole analysis in one place, with the RPN calculated as severity multiplied by occurrence multiplied by detection. Read the table as an argument, not as a result. Every number in it is invented for teaching.
| Ref | Functional failure | Failure mode | Local effect | System / end effect | Current control | S | O | D | RPN |
|---|---|---|---|---|---|---|---|---|---|
| M1 | FF1 low flow | Impeller and wear ring erosion | Head and flow decay | Cooling capacity degrades on peak days | BMS differential pressure trend, unreviewed | 6 | 5 | 4 | 120 |
| M2 | FF1 low flow | Discharge valve partially closed after work | Throttled duty point | Capacity shortfall, possible pump damage | BMS flow and pressure, work order correlation | 5 | 4 | 3 | 60 |
| M3 | FF1 then FF2 | Suction strainer blockage | Starved suction, cavitation | Progresses to total loss of duty | Monthly operator round, no alarm | 7 | 4 | 4 | 112 |
| M4 | FF2 no flow | Motor winding insulation breakdown | Motor fails, drive lost | Lead pump lost, long lead time to replace | None specific to windings | 8 | 2 | 6 | 96 |
| M5 | FF2 no flow | Coupling element fracture | Drive disconnected, motor runs free | Lead pump lost, reliance on changeover | Six-monthly PM, guard not removed | 8 | 3 | 5 | 120 |
| M6 | FF3 leakage | Mechanical seal face wear | Water at the seal, dripping | Plant room housekeeping and corrosion | Monthly visual round | 4 | 6 | 3 | 72 |
| M7 | FF3 leakage | Casing joint gasket relaxation | Weeping casing joint | Escalating leak, insulation damage | Monthly visual round | 5 | 3 | 4 | 60 |
| M8 | FF4 no start | Changeover interposing relay failure | No effect until called upon | Standby does not start, cooling lost | Changeover test recorded as a tick only | 7 | 3 | 8 | 168 |
| M9 | FF4 no start | Starter contactor contact welding | Fails to make or break cleanly | Start failure or uncontrolled running | Annual panel thermography | 7 | 3 | 6 | 126 |
| M10 | FF5 out of limits | Bearing degradation, lubrication starvation | Vibration and temperature rise | Progresses to seizure and secondary damage | Monthly round, no vibration route | 6 | 5 | 4 | 120 |
| M11 | FF5 out of limits | Grout and holding-down deterioration | Misalignment, elevated vibration | Accelerated bearing, seal and coupling wear | Not inspected | 5 | 3 | 5 | 75 |
Sorted by RPN: M8 at 168, M9 at 126, a three-way tie at 120 between M1, M5 and M10, then M3 at 112, M4 at 96, M11 at 75, M6 at 72, and M2 and M7 at 60. That ordering is useful. It is also, in two specific places, actively misleading.
11. What the RPN gets wrong, with these same numbers
The risk priority number is the reason most people search for an FMEA example, so it is worth calculating properly and then being honest about. Three defects are visible in my own worksheet above.
First, equal RPN does not mean equal risk. M1 and M5 both score 120. M1 is a gradual, non-safety, service-degrading mode with a trendable indicator. M5 is a sudden total loss of drive with a severity of 8 and no warning. Treating them as equivalent items on a ranked list is indefensible, and if the only thing on your report is the RPN column, that is exactly what you have done.
Second, a threshold hides high-severity rows. If a team drew a line and worked only the rows above 120, M4 would fall out of scope at 96. M4 is the highest-severity mode in the analysis, it has no detection control at all, and its consequence is the longest outage on the list. It scores low only because occurrence is 2. A rare event with a severe consequence and no detection is precisely the kind of risk you are meant to be looking for, and the arithmetic buries it.
Third, multiplying ordinal ranks is not arithmetic. Severity, occurrence and detection are ordered labels, not measured quantities. A severity of 8 is not twice a severity of 4; it is a different category of consequence, and the interval between a 7 and an 8 is not the same as between a 2 and a 3. Multiplying three such rankings produces a number with no units and no meaningful ratio properties, and the product space is lumpy: some values are reachable by many combinations, others by almost none. The same defect appears in a risk matrix built by multiplying likelihood and consequence bands, discussed in the criticality matrix article. It is worth understanding properly, because it changes how you present results to people making budget decisions.
| Comparison | Rows | S x O x D | RPN | What the RPN implies | What is actually true |
|---|---|---|---|---|---|
| Equal RPN, unequal risk | M1 impeller erosion M5 coupling fracture |
6 x 5 x 4 8 x 3 x 5 |
120 120 |
Identical priority, interchangeable on the ranked list | M1 is gradual, trendable and service-degrading. M5 is sudden total loss of duty at severity 8. They need different responses, not the same queue position. |
| High severity hidden below a cut | M4 winding breakdown (vs an illustrative cut at 120) |
8 x 2 x 6 | 96 | Below the line, no action required | Highest severity in the analysis, no detection control whatsoever, longest outage. Low only because occurrence is 2. This is the row a severity gate is designed to catch. |
| Ordinal product has no units | M8 at 168 vs M9 at 126 | 7 x 3 x 8 7 x 3 x 6 |
168 126 |
M8 is about a third more risky than M9 | The only real difference is one step of detection ranking. The 42 point gap is an artefact of multiplication, not a measured quantity of risk. |
On RPN thresholds
The 120 cut used in the table above is a number I invented to demonstrate a failure of thresholds. It is not a standard, not a convention, and not a recommendation. I am not aware of any authoritative source that sets a universal RPN action threshold, and any figure you see quoted as an industry standard cut-off should be treated as someone's local convention at best. If your organisation uses a threshold, it should be your own, documented, and never the sole trigger for action.
12. Better than a threshold: a severity gate and Action Priority
Knowing the RPN is weak is not an argument for abandoning prioritisation, only for prioritising in a defensible order. Two approaches do most of the work.
A severity gate applied before any RPN consideration. The simplest and most valuable change most teams can make. Before you look at a single RPN, take every row above a defined severity and treat it as in scope regardless of its product. In my worksheet a gate at severity 8 pulls in M4 and M5 immediately, and M4 is the row the arithmetic was hiding. The logic: severity reflects consequence, consequence is what you cannot recover from, and a low occurrence estimate is the least reliable of the three judgements anyway, because most organisations lack failure frequency data good enough to defend it.
Action Priority. The AIAG and VDA FMEA Handbook, first edition, June 2019, replaced RPN with an Action Priority approach. Be precise about what that document is: an industry handbook, not an ISO or IEC or ANSI standard, published jointly by the Automotive Industry Action Group and the German automotive association, whose authority flows through automotive customer contracts rather than from any standards body. Within that context it is influential, and its central move is a good one regardless of sector. Instead of multiplying the three ranks, Action Priority evaluates their combination against a defined logic and assigns each row a priority band of high, medium or low, with severity dominating. A high-severity row cannot fall to low priority on the strength of a favourable occurrence estimate, which is exactly the defect in my M4 row.
You do not need to adopt automotive documentation to borrow the principle. Severity leads, occurrence and detection modulate, and the output is a small number of priority bands a manager can act on rather than a list of pseudo-precise integers that invites arguing about whether 126 really outranks 120. The RPN column is a conversation starter, not an output. Its job is to get the right rows onto the table so that people with knowledge can argue about them. The moment it becomes a mechanical trigger, the analysis has stopped being risk management. For ranking assets by risk rather than ranking failure modes, the equipment criticality analysis guide covers the layer that should sit above this work.
13. Step eight: deciding actions, then re-scoring for the residual
An FMEA that ends at the RPN has produced nothing. The output is a set of actions with owners and dates, and then a re-score showing what the risk looks like once they are done. Re-scoring is the part almost everybody skips, and the part that proves the analysis changed something. The rule that keeps it honest: only change the score the action actually addresses. An inspection improves detection. Removing a cause improves occurrence. Severity moves only with a change to the design, the redundancy or the consequence. Watch what stays fixed below.
| Ref | Initial S/O/D | RPN | Action taken | What it changes | Residual S/O/D | Residual RPN |
|---|---|---|---|---|---|---|
| M8 | 7 / 3 / 8 | 168 | Convert the changeover test into a witnessed functional test under load, with the measured start time recorded against the work order, plus relay coil monitoring to the BMS | Detection only. A hidden failure becomes a found failure. | 7 / 3 / 3 | 63 |
| M9 | 7 / 3 / 6 | 126 | Add contactor contact inspection and resistance check to the annual electrical task, and extend thermography to the starter enclosure | Detection only. | 7 / 3 / 3 | 63 |
| M4 | 8 / 2 / 6 | 96 | Introduce annual insulation resistance and polarisation index testing with results trended year on year | Detection only. Severity 8 stands, because the consequence of a failed motor has not changed. | 8 / 2 / 3 | 48 |
| M5 | 8 / 3 / 5 | 120 | Require guard removal and coupling element inspection in the six-monthly task, and add a laser alignment check after any coupling work | Detection, and occurrence via better alignment. | 8 / 2 / 2 | 32 |
| M1 | 6 / 5 / 4 | 120 | Raise the loop water quality issue with the party who owns it, propose side-stream filtration, and assign monthly ownership of the pump differential pressure trend | Occurrence via cause removal, detection via trend ownership. | 6 / 3 / 2 | 36 |
| M10 | 6 / 5 / 4 | 120 | Move to a specified lubricant with a defined quantity and interval, and add a monthly vibration reading with a recorded value rather than a tick | Occurrence via correct lubrication, detection via measured trend. | 6 / 3 / 2 | 36 |
| M3 | 7 / 4 / 4 | 112 | Fit differential pressure indication across the strainer and alarm it to the BMS | Detection only. | 7 / 4 / 2 | 56 |
| M11 | 5 / 3 / 5 | 75 | Add baseplate, grout and holding-down bolt condition to the annual inspection | Detection only. | 5 / 3 / 3 | 45 |
| M2 | 5 / 4 / 3 | 60 | Add a signed valve line-up verification step to the close-out of any intrusive pump work | Occurrence via error-proofing the task. | 5 / 2 / 3 | 30 |
| M6 | 4 / 6 / 3 | 72 | Accept the mode. Hold a seal kit in the storeroom and treat replacement as planned consumable work | Nothing. A documented decision to accept is a legitimate outcome. | 4 / 6 / 3 | 72 |
| M7 | 5 / 3 / 4 | 60 | No action beyond existing round | Nothing. | 5 / 3 / 4 | 60 |
Three things to notice. Severity never moved in any row, because no action changed the design or the consequence. Most of the reduction came from detection, the normal pattern in maintenance FMEA and the reason detection is worth scoring honestly. And two rows were deliberately accepted with no action, which is a real answer. An analysis where every row generates a task has not prioritised anything.
Those actions are mostly task and evidence changes, so they land in a maintenance plan and a work order structure. The re-scored analysis is only real once the tasks exist with the recorded values the detection score assumed, which is where this connects to preventive maintenance programme design and to failure coding: the coded history from those work orders is what lets you revisit the occurrence scores with evidence rather than opinion. Any reasonably capable maintenance system will hold the worksheet and the linked tasks, and a spreadsheet plus disciplined work orders beats a dedicated module nobody updates.
14. A blank FMEA template you can lift
Here is the worksheet stripped back to an empty shell. Copy it, add a header block recording the item, the boundary, the operating context, the date, the participants and the scale definitions you agreed, and keep the initial and residual scores side by side so the change is visible without hunting through revisions.
| Ref | Function and performance standard | Functional failure | Failure mode (mechanism) | Local effect | System / end effect | Current control | S | O | D | RPN | Severity gate? | Action, owner, date | Residual S/O/D | Residual RPN |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
The two columns people leave out and should not are the severity gate flag, which forces someone to look at high-consequence rows independently of their product, and the residual pair, without which nobody can tell whether the analysis achieved anything.
15. Where the published documents sit
Briefly, and without pretending any of them will do the judgement for you. The FMEA guide covers this landscape properly.
- IEC 60812:2018, third edition, "Failure modes and effects analysis (FMEA and FMECA)". This is the international standard, published by the IEC . Worth knowing that the title changed at edition three; earlier editions sat under a broader analysis techniques for system reliability title, and citing that older wording is a common giveaway that someone has not opened the current document.
- AIAG & VDA FMEA Handbook, first edition, June 2019, available through the AIAG . An industry handbook, not a standard. It is the source of the widely cited seven-step approach and of Action Priority replacing the RPN.
- SAE J1739_202101, published by SAE International , is an SAE Recommended Practice deliberately aligned with the AIAG and VDA approach. It is an alternative compatible method rather than a replacement for it, and neither document is withdrawn in favour of the other.
All three are paywalled, so I am not quoting clause text. If your work is contractual or regulated, buy the applicable document and read it rather than relying on a summary, including this one.
16. The mistakes I see in real worksheets
- Functions written as descriptions. "Pump water" cannot be failed. Without a performance standard there are no functional failures and no structure.
- Failure modes written as symptoms. "Bearing failure" collapses several mechanisms with different scores and tasks into one unactionable row.
- Detection scored aspirationally. Scoring the inspection you plan rather than the control you have inflates the start and makes the improvement invisible.
- Hidden failures omitted. Protective devices, standby equipment and changeover logic fail silently. If nothing visible happens when a mode occurs, it belongs in the worksheet with a high detection score.
- Severity adjusted to get under a threshold. The quiet way analyses become worthless.
- No boundary statement. Two participants analysing different machines, discovered at review.
- No re-score. The residual column is the evidence the work mattered. Without it you have a document, not a result.
- Confusing failure mode with cause. An FMEA identifies what fails and what follows. Finding out why a specific failure actually happened is a different exercise, and its techniques are in the root cause analysis guide. A causal investigation inside an FMEA workshop means the workshop has lost its thread.
One broader point about where this belongs. An FMEA on every asset is unaffordable and unnecessary. It earns its place on the assets criticality screening has already flagged, and it is the natural input to the task-selection logic of a structured reliability programme, treated properly in the RCM introduction. One careful FMEA on the asset that keeps hurting you is worth more than a shelf of thin ones.
The idea to walk away with
The value of an FMEA is in the reasoning, not the number at the end of the row. Work the sequence properly, item and boundary, functions with performance standards, functional failures, mechanisms, effects at both levels, the controls you actually have, and the scores become explicable to someone who was not in the room. Compute the RPN, because it usefully surfaces rows for discussion, but apply a severity gate before you look at it, treat the ranking as an agenda rather than an answer, and finish with actions, owners and a re-score.
Final thoughts
If you take one thing from this worked example, make it the discipline of writing down what you actually have rather than what you meant to have. Weak FMEAs are commonly weak in the same place: the detection column describes an intention. Scores that follow from an honest control inventory are less flattering and far more useful, because they point at the specific gaps you can close with a task change next month.
And to repeat what matters most here: CHW-P-01 does not exist, and neither do its numbers. Take the sequence, the blank worksheet and the severity gate, and build your own scales against your own consequences. An FMEA is a record of your team's judgement about your equipment. Borrowed figures are not judgement, they are decoration.
Disclosure
Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.
Running your first FMEA on a critical asset?
Independent advisory on scale design, workshop facilitation, translating the output into a maintenance plan, and building the failure history that lets you revisit occurrence scores with evidence. 22+ years across utilities, oil and gas, manufacturing, government and facility operations.
Book a conversationRelated reading: FMEA: the method and the standards landscape, Common failure modes and how to analyse them, Equipment criticality analysis, Criticality matrix: ranking equipment by risk, Root cause analysis methods, From RUL to FMEA, RCM introduction, Failure codes: Problem, Cause, Action, Preventive maintenance: the complete guide.
Muhammad Abbas
CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.
Work with me