mail@mabbaz.com Abu Dhabi, UAE

Root Cause Analysis · Reliability · Worked Examples

Root Cause Analysis Examples: Equipment Failures

Most root cause analysis examples give you a symptom in one column and a root cause in the other, with nothing in between. That teaches nothing, because the whole skill sits in the reasoning you cannot see. This article does the opposite: seven illustrative equipment failures, invented for this article, each carried step by step from a vague symptom through the wrong turns to the systemic cause and the actions at every level.

Muhammad Abbas September 27, 2026 ~23 min read

A tidy root cause analysis example is almost useless as teaching material, and the tidiness is the problem. Real analysis is a sequence of hypotheses, most of them wrong, conducted on incomplete evidence by people who have other work waiting. What separates an analyst who finds something useful from one who writes "operator error" and closes the file is not knowing more methods. It is the habit of asking one more question after the answer already looks plausible. So this article shows the dead ends and the discarded hypotheses alongside the conclusions.

The message up front: every example below is my own invention, written to illustrate a reasoning pattern. None describes a real incident, a real organisation or any client work, and the handful of numbers in them are illustrative placeholders rather than measurements. Read them for the shape of the thinking, not for the figures. The pattern to take away is that the obvious answer is nearly always a component, and the useful answer is nearly always a decision.

1. What makes a worked example useful

This article stays deliberately inside one lane. It does not teach the RCA process, which belongs to the root cause analysis methods and step-by-step guide, and it does not help you choose between techniques, which belongs to root cause analysis tools: 8 methods explained. Read either of those for the method. Read this one to watch a method being used.

Three things make a worked example worth reading. The first is that it reaches more than one level. A finding that stops at the broken part is a repair note. A usable analysis separates the physical cause (what actually failed), the conditions that allowed it (the local circumstances that made that failure possible or likely), and the systemic factor (the decision, process, or absence of a process, that produced those conditions and will keep producing them elsewhere). Those are three findings with three different owners, and collapsing them into one sentence is how organisations end up fixing the same failure repeatedly.

The second is that it separates the cause of the failure from the cause of the escape. Something allowed this to develop unseen: a PM task that did not look for it, an inspection that looked but could not detect it, a trend nobody reviewed, or an operator report that went nowhere. The escape point is often the cheaper and more generalisable fix, because one detection improvement covers a class of failures while one component fix covers one asset.

The third is honesty about mess. A written example has hindsight: evidence arrives in a helpful order and the reasoning converges. Reality does not. The failed part went in a skip, the technician who did the previous repair has left, and the trend data stops where the historian was last reconfigured. Every example below includes at least one wrong turn, and even that understates how untidy it gets.

The test I apply to any finished analysis

If the stated root cause names a person or a component, and the corrective action is "retrain" or "replace", I have not got an analysis. I have got a repair with a paragraph attached. There is a real international standard for this discipline, IEC 62740:2015 "Root cause analysis (RCA)", and one of the more useful things it does is put the assigning of responsibility or liability outside the scope of RCA. Blame and analysis are different activities, and mixing them reliably ends the analysis early.

2. Example one: the pump bearing that keeps failing

Illustrative example. Reported symptom: "Pump noisy again, third bearing this year, please attend." No running hours, no note of which end, no vibration reading, no mention of the earlier repairs beyond a count somebody happens to remember.

The obvious answer: bad bearings, already written on the second work order. Inadequate for a simple reason: three failures in a year on one machine, while identical pumps on the same site run for years, is not a bearing quality story. Something in the installation is destroying bearings. A component that fails repeatedly in one location is almost never the cause of its own failure.

Evidence gathered: a thin work order history; the failed bearing, still in the workshop by luck; the wear pattern on the races; the alignment record, which did not exist; and a photograph from an unrelated job showing the suction pipework.

The reasoning, including the wrong turn: the first hypothesis was lubrication, because greasing intervals across that pump group were inconsistent. Reasonable, and wrong: the wear pattern was not consistent with lubricant starvation, and the same greasing regime applied to the pumps that were not failing. Discarding it mattered, because that answer was comfortable and would have closed the file. The second hypothesis was misalignment, and here the absence of evidence became the evidence. There was no alignment record because alignment had never been specified as a step, so nobody could say whether the coupling had ever been inside tolerance after any of the three repairs. The photograph suggested the mechanism: a suction pipe re-supported at some earlier date in a way that left a standing load on the casing. This is the linear chain the 5 Whys is built for.

Physical cause: bearing failure under a persistent off-design load. Conditions that allowed it: pipe strain transmitted into the casing, plus alignment never verified on reassembly. Systemic factor: the repair procedure for this pump class carried no alignment specification, no tolerance and no sign-off, so the quality of every reassembly depended on which fitter attended. Escape point: there was no post-repair verification of any kind, so a poor reassembly was indistinguishable from a good one until the next failure.

Actions: replace the bearing and correct the pipe support (weak, it fixes one machine); add an alignment step with a stated tolerance and a recorded reading to the reassembly procedure for the whole pump class, required before closure (strong, it changes every future repair); and take a post-repair vibration baseline so a poor reassembly shows up in days, a condition monitoring question rather than a repair one.

3. Example two: the chiller that trips on hot afternoons

Illustrative example. Reported symptom: "Chiller 2 tripped, reset and running, occupants complaining about temperature." Reported three times in one summer, each time late afternoon, each time reset successfully.

The obvious answer: a compressor or controls fault. Inadequate because the pattern, high ambient plus high load plus late afternoon, points at heat rejection. The compressor tripped on a protection device, and a protection device operating is usually the machine behaving correctly in response to something else.

Evidence gathered: the trip codes, the approach temperature across the condenser, the PM history, the actual condition of the coil, and a conversation with the technician whose name was on the PM completions.

The reasoning, including the wrong turn: refrigerant charge was checked first, because a low charge also produces heat-related trips. It was within specification and it went. Inspection then showed the condenser substantially fouled, which explained the physics at once, and that is where most investigations stop, with "clean the coil" as the action. The next question is the one that matters. The schedule contained a quarterly condenser clean and the records showed it complete every quarter. Both were true. The task was signed off without being performed, not out of dishonesty but because it fell when the unit could not be isolated, so it was marked complete and nothing in the system ever contradicted that. This one benefits from laying trips, ambient conditions and PM completions on one timeline.

Physical cause: a high condensing pressure trip from degraded heat rejection. Conditions: a fouled condenser at peak ambient with no spare capacity to absorb the loss. Systemic factor: PM completion was supervised by counting records rather than verifying outcomes, and the schedule was built without asking when the asset could actually be taken out of service. Escape point: approach temperature was on the panel and trending steadily worse for months with nobody looking, so the only detector of a failing PM regime was the failure itself.

Actions: clean the condenser (necessary, weak); reschedule into a window when isolation is possible and require an evidence artefact, a measured approach temperature before and after, rather than a tick (strong); and change how compliance is measured, since a figure that counts signatures measures paperwork rather than maintenance. That distinction is the subject of the PM KPIs and schedule compliance pillar.

4. Example three: the air handling unit belt that keeps going

Illustrative example. Reported symptom: "AHU 7 no airflow, belt snapped, replaced, back in service." Then the same again weeks later, and again. The unit had run for years on annual belt changes before this started.

The obvious answer: belt quality, tensioning, or a bearing beginning to drag. All plausible and all premature, because the analysis has not used its strongest clue. The unit used to work, so something changed at a point in time.

Evidence gathered: the failed belts; the part number actually fitted, from stores issue records rather than work order text; the equipment specification; the catalogue history; and the date of the first short-life failure.

The reasoning, including the wrong turn: tension was suspected first, and worth suspecting, since over-tensioning shortens belt life and no tension check existed in the task. It did not survive the timeline: practice had not changed, the same people did the work, and the failures began abruptly rather than drifting. Comparing the fitted part number against the specification found the difference. A visually similar belt of different construction had replaced the specified item in the catalogue, and it was not equivalent in the way that mattered. Nobody did anything wrong at the point of fitting; the fitter drew the part the system offered. This is what a formal change analysis handles, worth reaching for whenever "it used to be fine" appears in a report.

Physical cause: premature belt failure from a non-equivalent part. Conditions: the substitute catalogued against the asset as though equivalent, so the correct part was no longer obtainable. Systemic factor: a purchasing substitution made on commercial grounds without engineering review and without telling maintenance, so a specification decision was taken by a process with no visibility of the consequence. Escape point: the recurrence itself. Three short-life failures generated three work orders and no review, because nothing flags repeat failures.

Where this analysis nearly went wrong

The tensioning hypothesis was attractive because it had an owner, a technician, and a cheap action, a toolbox talk. It would have been signed off happily. Hypotheses pointing at the people nearest the failure are the easiest to accept and the hardest to disprove, which is why they deserve more scepticism than the ones pointing at a process nobody in the room controls.

5. Example four: the hunting control valve and the out-of-specification process

Illustrative example. Reported symptom: "Process temperature swinging, product out of specification intermittently, valve sounds like it is hunting."

The obvious answer: the valve, because it is the moving part making a noise everyone can hear. Inadequate because a valve that hunts is frequently a valve obediently following an unstable command. Replacing it here produces a new valve that hunts.

Evidence gathered: valve position and controller output logged together, the single most useful thing done; a stroke test; the tuning parameters as configured; the commissioning records; and the change log, which did not exist for this loop.

The reasoning, including the wrong turn: the first hypothesis was mechanical, stiction in the stem, a genuine and common cause of hunting. The stroke test made it unlikely, since position followed command cleanly both ways. The logged data settled it: the command itself was oscillating and the valve was tracking it faithfully. That moved the question to the controller, where the tuning was aggressive in a way nobody could account for. Current values did not match the commissioning records, there was no record of any change, and the people who might have known had moved on. The honest conclusion was that a setting had been altered during or after commissioning, plausibly to solve a temporary problem, and had then persisted as the configuration. Note what happened: the analysis established that an undocumented change existed without ever establishing who made it, and that was enough to act on.

Physical cause: an unstable control loop driving valve oscillation. Conditions: tuning inappropriate for current process conditions, persisting unnoticed because it produced tolerable behaviour most of the time. Systemic factor: no change control over control system configuration, so the live configuration was not the documented one and nobody could tell which parameters had drifted. Escape point: no comparison of live configuration against the commissioning baseline existed as a task.

Actions: retune the loop (necessary, and weak, because the next undocumented change does the same); establish a configuration baseline with periodic comparison against live values and a change record for every parameter change (strong); and treat configuration drift as a failure mode in its own right, the kind a failure mode review should surface before it causes a quality problem.

6. Example five: the standby generator that would not take load

Illustrative example. Reported symptom: "Mains failed, generator started, tripped shortly after the changeover, site lost supply." The report is almost apologetic, because the generator did start. That is the detail that makes this the most instructive example in the set, and it gets the fullest treatment here.

The obvious answer: the generator failed, so investigate the generator. That framing is wrong before it begins, because the failure under investigation is not "the generator stopped". It is "a protective function was unavailable when it was called on, and nobody knew". The asset has a hidden function. It sits idle, and its failure is not evident in normal operation. There is no complaint, no noise, no alarm, no drop in output. A hidden failure announces itself only when the protected event occurs, which is the worst possible moment to discover it.

Evidence gathered: the controller event log, recording a start, a changeover and a trip with a fault code; the fuel system, filters and water separator; the cooling system; the battery and charger; two years of test history; and, crucially, the actual written content of the monthly test task.

The reasoning, including two wrong turns. The first hypothesis was fuel, the statistically sensible first guess for a generator that starts and then fails. Fuel quality, filter condition and the water separator were all acceptable. The second was the battery and starting system, which also fails often, and it was eliminated partly because the unit had started without difficulty. Both were necessary eliminations rather than wasted work, but neither addressed the shape of the failure. The unit ran and then failed when load was applied, and the only difference between the successful monthly test and the real event was load.

That reframing is where the analysis turned. The monthly test task, read carefully, said in substance: start the generator, run it for a period, confirm it runs, stop it, record complete. It had been performed faithfully, on time, every month, for years, by conscientious people. And it proved almost nothing about the only capability that mattered. A no-load run demonstrates that the engine starts and idles. It does not demonstrate that the cooling system can reject heat at load, that fuel delivery holds up at load, that the alternator and its regulation are sound under load, that the changeover sequence completes, or that the site's actual load is within what the machine can carry now as opposed to when it was commissioned. This was a test regime with a high compliance rate that tested the wrong thing, which is more dangerous than no test at all, because it manufactures confidence.

Physical cause: a protection trip on load caused by a degraded system that idling did not exercise; which system matters less than the category, and in a real investigation the fault code and strip-down would resolve it. Conditions: years of light no-load running, itself a degradation mechanism on a diesel rather than a neutral activity, combined with a site load that had grown since commissioning without anyone revisiting the rating. Systemic factor: a test regime designed to demonstrate starting rather than the function the asset exists to provide, on an asset whose maintenance plan had never been derived from the consequence of its failure. Escape point: the heart of it. Not a missed inspection, but a test that could not detect the failure it existed to detect, performed diligently and recorded as successful every month.

Hidden failures need a different kind of task

A protective device whose failure is not evident in normal operation needs a task whose purpose is to find out whether the function still works, not to service the device. That distinction, and the consequence logic that leads to it, is central to how reliability-centred maintenance handles hidden functions; SAE JA1011_202411 sets out the criteria an analysis must satisfy to be called RCM, though it deliberately does not prescribe a process. The practical version is a blunt question: for every asset whose job is to respond to an event that has not happened yet, what task proves it still can, and when did that task last actually prove it? Start with the RCM introduction if this is new territory.

Actions: repair the defect found (weak, and it will pass the old test afterwards regardless); replace the no-load run with a load test against a realistic load, with recorded parameters, plus a periodic review of site load against machine rating (strong, it turns a meaningless test into a real one); and extend the same question to every other hidden-function asset on site, fire pumps, emergency lighting, trip systems, changeover panels, because if the reasoning applies to the generator it almost certainly applies to several of them. That last one is the strongest action anywhere in this article, because it acts on a class rather than an asset. A fault tree suits this kind of on-demand protective system, reasoning from an undesired top event down through the combinations that produce it; the reference is IEC 61025:2006, still the current edition. Whether the resulting change is logged as a correction or a preventive action is a CAPA question rather than an analysis one.

7. Example six: the water leak that keeps causing damage

Illustrative example. Reported symptom: "Water damage to ceiling in the same area again, third time, please investigate." Each previous occurrence had its own work order, and each was closed as complete.

The obvious answer: a pipe in poor condition that needs replacing, which may well be true and is still not an analysis. The interesting question is not why the pipe leaks. It is why three closed work orders produced no durable outcome.

Evidence gathered: the three work orders and their closure text; the physical repair, inspected rather than assumed; the parts issued against each job; and the reporting history, including complaints logged through a channel that generated no work orders at all.

The reasoning, including the wrong turn: the first hypothesis was systemic pipe deterioration, which would have justified a survey and a capital case. Inspecting the repair itself redirected everything. What was in place was a temporary containment, competently applied as a stop-gap by someone working at night with the materials to hand and every intention of coming back. The work order was closed as complete because "complete" was the only state that cleared it from the list. There was no way to record "made safe, permanent repair required", so the fact that a temporary repair existed was lost at the moment of closure, and had been lost twice more since.

Physical cause: failure of a temporary repair, as temporary repairs are designed to do. Conditions: that repair left in service indefinitely because nobody downstream knew it was temporary. Systemic factor: a closure practice with no state between open and complete and no mechanism to raise a follow-on job, so the system structurally could not carry forward the knowledge that work remained. Escape point: nothing reviewed closed work orders for repeat locations, and complaints arriving outside the work order channel never entered the history, so the recurrence was invisible in the data while being obvious to the occupants.

Actions: execute the permanent repair (necessary, weak); introduce a closure state for temporary work that mandatorily generates a linked follow-on job with an owner and a date, so the backlog carries the truth rather than hiding it (strong); and route all reporting channels into one history so recurrence is detectable. Once temporary repairs are visible they become a measurable category, which connects to how backlog is tracked and to the discipline in the corrective maintenance guide. A backlog that looks small because temporary repairs are recorded as complete is not a small backlog. It is an undocumented one.

8. Example seven: the mundane fault that keeps coming back

Illustrative example. Reported symptom: "Lights out in corridor, again." Or a small pump, or a door closer. Deliberately unglamorous, because the method applies well below the level of dramatic failures, and this is where most maintenance effort actually goes.

The obvious answer: nothing worth analysing. A lamp failed, someone changed it, move on. Inadequate because the aggregate is not trivial even when each instance is. The reason such a fault recurs is usually not that anybody misdiagnosed it. It is that nobody has ever been able to see that it recurs.

Evidence gathered, or rather attempted: pulling the history for the location, which is where the analysis met its real obstacle. The jobs existed but were spread across several free-text descriptions of the same area, some raised against a generic building asset and some against no asset at all, with failure coding blank or defaulted. No single query could show the pattern.

The reasoning, including the wrong turn: the first hypothesis was a supply quality or fitting-type problem, arrived at by standing in the corridor and reasoning about it, which is a respectable way to generate a hypothesis and a poor way to test one. Testing it required history, and the attempt to get the history was the actual finding. Reconstructing the record by hand showed a genuine recurrence pattern, and once visible it pointed at a specific local condition, plausibly a control or environmental factor affecting that run of fittings. But most of the effort went on data archaeology rather than on the failure, and that is the transferable lesson. The technique that matters here is not a causal one at all. It is Pareto analysis over the work history, and it only works if the history can be grouped.

Physical cause: a local condition shortening component life, identifiable only once the pattern was visible. Conditions: repeated like-for-like replacement with no diagnosis, because each attendance looked like a first occurrence. Systemic factor: asset registration and failure coding practice made recurrence undetectable, so nothing could notice that the same job was being paid for repeatedly. Escape point: the data itself. This failure did not escape an inspection, it escaped the record.

Actions: fix the local condition (proportionate, and only possible once the pattern was visible); register the assets properly and enforce structured failure coding so recurrence surfaces without manual archaeology (strong, and it pays back across every asset class); and run a standing repeat-failure report, the cheapest reliability improvement available to most organisations because it needs no new data, only usable coding. The structure that makes this work is the subject of the failure codes pillar. For the taxonomy underneath it, ISO 14224:2016 is the reference standard for collecting and exchanging reliability and maintenance data; OREDA, often mentioned alongside it, is a proprietary database rather than a standard.

9. All seven examples, side by side

Laid out together, the same structure repeats. The reported symptom is vague, the obvious answer is a component, and the useful finding is a decision somebody made, often years earlier and usually for a defensible reason at the time. Every row below is illustrative and invented.

Reported symptom Obvious answer Physical cause Systemic factor Escape point
Pump noisy, third bearing this year Bad bearings Bearing failure under off-design load No alignment specification or sign-off in the repair procedure No post-repair verification of any kind
Chiller tripped, reset, occupants complaining Compressor or controls High condensing pressure trip, degraded heat rejection PM completion supervised by counting records, not verifying outcomes Worsening approach temperature visible and unreviewed
AHU belt snapped again Belt quality or tension Premature failure of a non-equivalent part Purchasing substitution with no engineering review or notification No repeat-failure flag; three jobs, no review
Temperature swinging, valve hunting The valve Unstable control loop driving valve oscillation No change control over control system configuration Live configuration never compared to baseline
Generator started then tripped on changeover The generator failed Protection trip on load, system not exercised by idling Test regime designed to prove starting, not function A diligent monthly test that could not detect the failure
Water damage in the same area again Pipe needs replacing Failure of a temporary repair Closure practice with no state between open and complete No repeat-location review; reports outside the work order channel
Lights out in corridor, again Nothing worth analysing Local condition shortening component life Asset registration and failure coding make recurrence invisible The record itself; each attendance looked like a first

10. The patterns visible across the set

Five patterns recur strongly enough across the set to be worth treating as working assumptions on your next investigation.

The obvious answer is a component and the useful answer is a decision. In every example, the finding that generalises is a choice somebody made: what to specify, what to buy, what to test, how to close a job, how to code a fault. Components are where failures become visible, not where they originate.

A large share of equipment failures trace back to a previous intervention rather than to wear. The bearing, the belt and the valve all failed because of something done to them, not because they ran out of life. That makes maintenance activity itself a failure source, and it argues for the things organisations skip under pressure: post-work verification, baselines after intervention, specifications for reassembly. "What was the last thing anyone did to this asset?" earns its place as a standing first question.

Hidden failures are systematically under-detected because the test proves the wrong thing. Any asset that waits for an event has this exposure, and it is worst where compliance is highest, because a well-run programme of the wrong test produces confident, well-documented, entirely unfounded assurance.

Temporary repairs recorded as permanent are a major source of recurrence. Almost always a system design problem rather than a behaviour problem. If the only closure state available is "complete", people will use it, and the organisation loses the most important fact about the work. Look at what states your closure workflow offers before concluding anyone is cutting corners.

"The part was bad" almost always dissolves on examination into a specification, storage, selection or substitution question. Sometimes parts genuinely are defective, and establishing that properly is the territory of physical examination, covered in failure analysis methods. It should be a conclusion you earn, not the first landing place.

What these patterns are not

These are observations from advisory work, offered as hypotheses worth testing on your own failure history, not as measured proportions. I have deliberately published no figure for how often failures trace to a previous intervention, or how much recurrence temporary repairs cause, because no credible universal figure exists and any number I offered would be invented. Run the analysis on your own data and you will get your own distribution, which is the only one that should drive your decisions.

11. Weak and strong actions, drawn from the examples

Judge an analysis by what it produced. Weak actions depend on a person remembering, noticing or being careful. Strong actions change the system so the failure becomes harder regardless of who is on shift. Both have their place, since the weak action is often what gets the asset running tonight, but a set containing only weak ones means the analysis never reached a systemic cause.

Example Weak action (necessary, not sufficient) Strong action (changes the system)
Pump bearing Replace bearing, correct pipe support Alignment tolerance in the procedure, recorded reading required to close
Chiller trip Clean the condenser; remind the team Reschedule to a feasible window; require a measured before and after; measure PM by outcome
AHU belt Fit another belt; toolbox talk on tensioning Engineering review gate on part substitutions; automatic repeat-failure flag
Control valve Retune the loop; replace the valve Configuration baseline, periodic comparison, change record for every parameter
Standby generator Repair the defect; repeat the monthly run Load test with recorded parameters; load-versus-rating review; same question applied to every hidden-function asset
Water leak Do the permanent repair this time Temporary closure state that mandatorily raises a linked follow-on job; single reporting channel
Recurring small fault Change the component again Proper asset registration and structured failure coding; standing repeat-failure report

Notice how few of those right-hand actions are engineering work. Most are changes to a procedure, a purchasing gate, a configuration or a report: cheap in capital, expensive in attention, which is precisely why they get dropped once the immediate failure is cleared and the next one is waiting.

12. How these analyses go wrong in practice

Written examples make the reasoning look inevitable. In practice analyses stop early, and they stop in four recognisable places.

Stopping at the component. The easiest to spot, because the corrective action is a part number. Ask whether the finding could recur elsewhere on site; if it could, and your action does not touch that other place, you stopped too early.

Stopping at the technician. Somebody did something wrong, so the action is retraining. Occasionally right, far more often not: as in the chiller and the water leak, the person was responding sensibly to a system that offered no better option, and retraining changes nothing because the next person faces the same constraint. IEC 62740:2015 places the assigning of responsibility or liability outside the scope of RCA, which is method rather than softness, because an analysis aimed at establishing fault stops the moment fault is established, always before the useful finding. Where accountability must be handled alongside causation, that is a distinct discipline covered in the incident investigation guide.

Stopping when a plausible story fits the facts. The subtlest one, and it caught several examples above. A story that explains the evidence is not the same as a story the evidence requires. Lubrication explained the bearing, tension explained the belt, stiction explained the valve; all coherent, all wrong. Ask what else would be true if the hypothesis held, then go and look.

Stopping where the next question would cost money. The least discussed and probably the most consequential. Analyses converge remarkably often on causes whose remedies are free, more care and better communication, and far less often on causes needing a capital case, a purchasing policy change, or an admission that a commissioning decision was wrong. That is rarely conscious. It is the analysis finding the nearest comfortable exit, and the only real defence is for the review to be led by someone who does not own the budget the answer might land on.

The honest limitation: not every failure deserves this

Each example above represents real hours of skilled time, and that time is scarce. Applied to everything, this depth of analysis collapses under its own weight and the organisation quietly abandons it. Criticality decides which failures get it: consequence of failure, recurrence, and whether the finding would generalise beyond the one asset. Everything else gets a competent repair, a correct failure code, and a place in the data so that if it does become a pattern, the pattern is visible. That is the point of the criticality analysis work: it tells you where to spend the analytical effort.

A closing note on methods, since this article has named several without teaching any. Choosing between them belongs in the tools comparison, but the version visible across these seven examples is that the technique follows the shape of the problem. A linear chain wants a sequence of questions. Competing possibilities want a structured cause grouping of the sort a fishbone diagram provides. Something that used to work wants a change analysis. An on-demand protective function wants a fault tree, per IEC 61025:2006. Recurring small faults want Pareto and clean coding under a taxonomy such as ISO 14224:2016. Working forward from a design to anticipate failures is FMEA instead, currently IEC 60812:2018, and the broader catalogue of assessment techniques is IEC 31010:2019, noting the designation carefully: IEC 31010, not ISO 31010. All voluntary, mostly paywalled, and none will conduct an analysis for you.

The idea to walk away with

Worked examples matter more than method descriptions because the method is easy and the discipline is hard. Anyone can learn to ask why five times in an afternoon. The difficulty is asking the sixth question when the fifth answer is already acceptable to everyone in the room, and looking for the evidence that would kill your own hypothesis rather than the evidence that supports it.

Across all seven illustrative failures the same shape holds. A physical cause tells you what to repair. Conditions tell you what to change locally. A systemic factor tells you what to change so the failure stops appearing on assets you have not looked at yet. Separately, the escape point tells you how to find the next one earlier. An analysis producing all four is worth the hours it cost; one producing only the first is a repair note, however long the document.

Final thoughts

One habit to take away: on your next recurring failure, before you touch the asset, write down the obvious answer and treat it as the hypothesis most likely to be wrong. Then ask what the last intervention on that asset was, what should have caught this and why it did not, and whether the finding could recur elsewhere. Those four questions do most of the work of a formal method, and they are the ones the examples above turned on.

Be realistic about depth, too. Most organisations do not need more RCA. They need a few properly finished analyses on genuinely important failures, plus clean enough data that mundane recurrences surface without manual reconstruction. The generator is the example I would leave you with, because it generalises furthest: somewhere on your site sits an asset waiting for an event, with a test regime that has a perfect completion record and proves nothing about whether it will work. Finding that before the event does is the highest-value analysis available to you, and it needs no new sensors.

Disclosure

Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.

Failures that keep coming back?

Independent advisory on repeat-failure analysis, failure coding that makes recurrence visible, hidden-function test regimes, and the work order practices that let temporary repairs disappear. 22+ years across utilities, oil and gas, manufacturing, government and facility operations.

Book a conversation

Related reading: RCA methods and step-by-step guide, RCA tools: 8 methods explained, Failure codes: Problem, Cause, Action, RCM introduction, CAPA explained, Equipment criticality analysis. Standards bodies referenced: IEC , ISO , SAE International .

Muhammad Abbas

CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.

Work with me
MAbbaz.com
© MAbbaz.com