mail@mabbaz.com Abu Dhabi, UAE

Root Cause Analysis · Reliability · Investigation

Root Cause Analysis: Methods and a Step-by-Step Guide

Root cause analysis is the discipline of working backwards from something that went wrong to the conditions that allowed it, and stopping only when you reach a cause you can actually do something about. This is the full guide: what RCA means, the process step by step, the methods landscape, the evidence discipline almost every training course skips, and the honest failure modes that turn a good investigation into paperwork.

Muhammad Abbas September 27, 2026 ~26 min read

I have read a great many completed root cause analysis reports over twenty-two years of CMMS, EAM and CAFM work, and most share one defect. They describe the failure in more detail than anybody needed, name a cause that is really the failure restated, and close with a corrective action amounting to a promise to be more careful. Nobody sets out to do that. The investigators were competent and the equipment genuinely broke. What went missing was the discipline: a precise problem statement, evidence gathered before it evaporated, a timeline, candidate causes tested rather than assumed, and corrective actions that change a condition rather than describe an intention. This guide is about that discipline.

The message up front: the only distinction in root cause analysis that really matters is the difference between a cause you can act on and a cause you can only describe. "The bearing failed" describes. "Bearings on this pump class are greased to no defined quantity because the task instruction never specified one" can be acted on. Everything else in RCA, every method, every diagram, every template, exists to move you from the first kind of statement to the second. If your report ends on a describing cause, the investigation is not finished, however thorough it looks.

1. What root cause analysis actually is

Root cause analysis is a structured, after-the-event investigation that identifies the causes of an undesired outcome so that action can be taken to stop it recurring. Three parts of that sentence carry weight. It is structured, which separates it from an experienced person's hunch, however often the hunch turns out to be right. It is after the event, which separates it from FMEA and other predictive techniques that ask what could go wrong before anything has. And its purpose is prevention of recurrence, not explanation, and certainly not attribution of blame.

There is a real international standard here, which surprises people told RCA is purely a folk practice. IEC 62740:2015 "Root cause analysis (RCA)" , adopted in Europe as EN 62740:2015, sets out the principles of RCA, specifies the steps an RCA process should include, and in its annex describes recognised techniques including the "Why" method that everyone calls 5 Whys, the Fishbone or Ishikawa diagram, events and causal factors charting, causes trees and fault trees. Two scoping decisions are worth knowing. It addresses only analysis after an event has occurred, so it is not a risk-assessment document. And it explicitly excludes the assignment of responsibility or liability: the standard defining the discipline says in its own scope that apportioning blame is not part of it.

The standard is paywalled, as almost all of them are, so I will not quote clause wording. What matters for a practitioner is the shape it confirms: RCA is a process with definable steps, the well-known tools are recognised techniques within that process rather than the process itself, and the output is intended to drive corrective action, not to produce a verdict on a person.

One more framing point, because it is easy to reach for the wrong technique. FMEA and fault tree analysis are mainly forward-looking: they start from a system and ask what failures are possible and what their effects or combinations would be. RCA starts from something that has already happened and works backwards. The techniques overlap, and a fault tree built during design can be reused in an investigation, but the direction of travel differs. For the forward-looking side of that pairing, the RUL to FMEA guide on modelling equipment failure covers it, and the dedicated FMEA guide goes deeper still.

2. The distinction that matters most: causes you can act on

Most weak RCA is weak not because the investigator lacked a method, but because nobody applied a test to the candidate cause before accepting it. The test is a single question: if I remove or change this condition, does the event become impossible or substantially less likely, and is removing or changing it within somebody's control? A statement passing both halves is a cause worth acting on. One failing either half is a description.

"The bearing failed" fails on both halves: you cannot remove bearing failure as a condition, and there is nothing to control. "Component failed" is the same sentence with the noun removed. Now consider "the bearing ran without grease because the lubrication task was written with no grease quantity, so technicians applied whatever seemed right and one applied almost nothing". Specify quantity and grease type on the task and the failure mode becomes much less likely. Somebody owns the task library. It passes.

The same logic disposes of "operator error", the most common stopping point in industrial investigations and almost never a root cause. It describes what a person did and offers that as though it explained why. Accept it and you have closed the investigation exactly where it was about to become useful. The questions worth asking instead:

  • Was the correct action obvious? If two valves are identical, adjacent, unlabelled and operated in opposite senses, selecting the wrong one is a design outcome, not a personal failing.
  • Was the person competent for this specific task, not just generally experienced? Competence is task-shaped, and induction records are not competence records.
  • Did the procedure match reality? Procedures that cannot be followed as written get worked around, and the workaround becomes the real procedure long before anyone documents it.
  • Was there time and staffing to do it properly? A forty minute task scheduled in a twenty minute window will be done incorrectly by somebody eventually.
  • Was feedback available? Could the person tell from the equipment or instrumentation that the action had not had the intended effect, in time to correct it?
  • Had this happened before, quietly? Near misses absorbed without record are the strongest available signal that the condition is systemic.

Answer those and "operator error" dissolves into something fixable: a labelling standard, a revised task instruction, a realistic duration in the planning system, an interlock, a competency assessment. That is the whole move.

The test to apply to every candidate cause

If I remove or change this condition, does the event become impossible or substantially less likely, and is it within somebody's control to remove or change? Both halves must be yes. A cause that fails either half is a description of the failure, not a cause of it, and an action written against it will not prevent recurrence.

3. The RCA process, step by step

The sequence below is the one I would run, expressed in the language of a maintenance or facilities operation rather than standards prose. The steps are ordered for a reason: serious RCA failures commonly come from taking them out of order, usually by leaping to causes before the evidence was secured or the timeline built.

Step 1: define the problem precisely. A vague problem statement guarantees a vague conclusion. "Pump failure" is not a problem statement. A usable one fixes the object, the deviation, the time, the place and the consequence: "Chilled water pump CHWP-03 tripped on motor overload at 02:14 on 4 March, for the third time in nine weeks, causing loss of cooling to Block C for 3 hours 40 minutes." That sentence tells you what to investigate, that recurrence is part of the problem, and what the consequence was, which later tells you how much corrective action is proportionate.

Step 2: preserve and gather evidence. This step belongs here, second, because evidence is perishable and it is disappearing while the investigation is being scheduled. It gets its own section below because it is the part most RCA training skips and the part that most often determines whether the analysis can reach a conclusion at all.

Step 3: build the timeline. Lay the sequence out in order with times attached: normal operation, the first deviation, alarms, operator actions, the failure, the response. A timeline does three things discussion does not. It exposes gaps, which is usually where the interesting part is. It kills impossible theories immediately, because a cause cannot postdate its effect. And it surfaces the standing conditions present in parallel, which a purely sequential account hides. Build it from records: alarm and event logs, SCADA or BMS trends, historians, access-control records, and work order history in the CMMS.

Step 4: identify causal factors. Working from the timeline, list everything that contributed, without yet ranking or filtering: the direct physical mechanism, the conditions that allowed it, the controls that should have caught it and did not, and the organisational conditions behind those. Being generous here is correct; premature narrowing is the commonest way to miss the cause that mattered.

Step 5: test candidate causes against the evidence. This is the step most often skipped, and the one that separates analysis from storytelling. For each candidate, ask what the evidence would look like if it were true, then check. If lost lubrication is the candidate, the bearing surfaces should show a particular thermal and wear signature and the lubrication records should support it. If they do not, the candidate is wrong regardless of how plausible it sounded in the meeting. Actively look for the evidence that would disconfirm your favourite theory. A cause that survives a genuine attempt to disprove it is worth acting on; one that was never challenged is not.

Step 6: identify the causes worth acting on. You will usually end with several validated causes at different depths, and you do not act on all equally. The practical rule: stop when you reach a cause whose correction is within the organisation's control and proportionate to the consequence, and where going further only produces conditions you cannot influence.

Step 7: define corrective actions. Each action must change a condition, have a single named owner, have a date, and be verifiable by somebody who was not involved in the investigation. Section 8 deals with how to tell a real action from a restatement.

Step 8: verify implementation, then confirm effectiveness. These are two separate checks and conflating them is a standing weakness in most RCA programmes. Verifying implementation asks whether the action was done: the instruction was revised, the guard was fitted, the training was delivered. Confirming effectiveness asks whether the event stopped recurring, which needs a defined measure and a defined review point, typically one to three maintenance cycles later depending on the failure interval. A programme that only ever verifies implementation will keep closing out actions on failures that are still happening.

4. Evidence discipline: the part most RCA training skips

Nearly every RCA course spends its time on the analytical tools and almost none on evidence. This is backwards. The tools are straightforward and can be learned in an afternoon. Evidence is difficult, perishable, and once it is gone the most rigorous 5 Whys in the world is a structured guess. If I could change one thing about how organisations run RCA, it would be to move evidence preservation from an afterthought to the first operational response once the situation is safe. There are four categories, and they degrade at very different speeds.

  • Physical evidence is the scene and its condition: valve and switch positions, isolation status, guarding, leaks, debris, temporary arrangements, what was where. It degrades within minutes to hours, because the first thing any competent operation does after a failure is tidy up and restore service. Photograph widely before anything is touched, including what looks irrelevant, record valve and breaker positions before anyone resets them, and date and orient the photographs, because one nobody can place is nearly worthless six weeks later.
  • Parts and samples survive indefinitely if retained and are lost permanently if they are not. The failed bearing, the ruptured hose, the burnt contactor, an oil sample taken before the system is flushed: all are the primary record of the physical mechanism, and all routinely leave site in a scrap skip within a day. My standing advice is blunt: nothing failed leaves site until the investigation says so, and the store needs somewhere to put it. Where the mechanism is metallurgical or chemical, laboratory examination of the retained part is often the only way to distinguish between candidate causes, and you cannot commission it against a part you threw away.
  • Data is alarm and event logs, SCADA and BMS trends, historian records, protection relay and drive fault records, access and CCTV records, and CMMS history. Data feels safe because it is digital, and it is not. Many controllers hold a fixed-depth buffer and overwrite the oldest records. High-resolution trend data is commonly aggregated to a coarser interval after a short retention period, which is precisely the resolution you needed, and CCTV retention is often days. Export the raw data on day one, at full resolution, over a wider window either side of the event than you think you need.
  • Witness accounts degrade fastest of all, and not by being forgotten. They degrade by being reconstructed: within hours, people have discussed the event, heard a theory, and unconsciously fitted their recollection to it. Interview individually and early, ask for the sequence in the person's own words before any specific questions, and separate what they observed from what they concluded. "I heard a change in the pump note about a minute before it tripped" is evidence. "I think the bearing had been going for a while" is a conclusion, worth recording as a lead, not as an observation.
The limitation nobody puts in the report

Where the evidence is gone, the honest thing is to say so in the report and state the resulting confidence level, rather than presenting a plausible mechanism as an established one. A report that says "the failed component was scrapped before retention, so the distinction between contamination and fatigue could not be established; the corrective actions address both" is far more useful, and far more credible, than one that picks a mechanism and sounds certain. Unacknowledged uncertainty is what makes an organisation stop trusting its own RCA output.

A practical footnote, and the one place software genuinely belongs in this discussion: if failure history is captured with consistent problem, cause and action coding, the recurrence question in step one is answerable in minutes rather than being a matter of who remembers what. If it is not, every investigation starts from zero. That is a data-discipline issue rather than a product issue, and the failure codes guide on Problem, Cause and Action covers how to structure it so the history stays usable. The same applies to downtime: if you cannot state the consequence in hours, you cannot judge what corrective action is proportionate, which the backlog and downtime tracking guide deals with. Readers without a system of record at all should start with the introduction to what a CMMS is.

5. The methods landscape: an overview

The methods are tools inside the process described in section 3, not alternatives to it, and choosing between them matters much less than running the process properly. What follows is enough to pick the right one and know its limits; the dedicated guides carry the detail.

Method What it is good for Where it breaks down Read more
5 Whys (the "Why" method) Fast, low-ceremony drilling on a single clear causal chain. Excellent for shop-floor use and for teaching the habit of not stopping at the first answer. Assumes one linear chain, so it hides the common case of several contributing conditions. Very sensitive to the investigator's assumptions: two people get two different chains from the same failure. 5 Whys guide
Ishikawa / fishbone diagram Structured group brainstorming across categories so that whole classes of cause are not forgotten. Good for widening the search before narrowing it. Generates possibilities, not conclusions. Produces a crowded diagram that feels like an answer while nothing on it has been tested against evidence. Fishbone guide
Fault tree analysis Logical decomposition of an undesired event through AND and OR gates. Handles combinations of conditions and supports quantification where failure-rate data exists. Effort-heavy and needs system knowledge. Overkill for routine equipment failures, and quantification is only as good as the input data, which is usually weaker than it looks. Fault tree guide
Pareto analysis Deciding what to investigate. Ranks failure categories by frequency or cost so effort concentrates where the loss is. Not an RCA method at all: it selects targets, it does not find causes. Also biased by whatever your coding scheme happens to record. Pareto guide
Events and causal factors charting Complex, multi-party incidents unfolding over time. Makes the sequence, the parallel conditions and the gaps visible in one picture. Time-consuming to build and maintain, and disproportionate for a single-component failure. Needs reasonably good time-stamped records to be worth doing. Incident investigation guide
8D A full problem-solving container with containment, verification and closure built in. Common where a customer requires a structured response to a defect. It is a framework wrapped around an RCA, not an analysis technique: you still need a method for the cause step. Heavy for internal use unless a customer requires the format. 8D guide

A note on standing in the standards, since it is misrepresented in both directions. No standard prescribes how to run a 5 Whys or draw a fishbone, but IEC 62740:2015 describes both as recognised RCA techniques, which is a stronger position than the "informal tool" label they usually get. Fault tree analysis has its own standard, IEC 61025:2006, still the current edition. FMEA and FMECA are covered by IEC 60812:2018. Several of these techniques also appear in IEC 31010:2019 "Risk management - Risk assessment techniques", worth naming correctly: it is IEC 31010, not ISO 31010. Bowtie analysis, sometimes reached for in this space, is not a standard at all; it is described in a 2018 concept book published jointly by the CCPS and the Energy Institute, and "the bowtie standard" is simply wrong. All of these are voluntary documents; none is law anywhere by itself.

For a side-by-side treatment of a wider set of techniques, see the RCA tools guide covering eight methods. For the physical and metallurgical side, how you establish what the component actually did, the failure analysis methods guide is the companion piece, and for more fully worked cases the RCA examples from equipment failures collection carries them.

6. Two worked examples

Both examples below are hypothetical and illustrative. I have deliberately not used client situations, and the equipment, intervals and figures are invented to make the reasoning visible.

Example A: the chilled water pump that kept tripping. The problem statement from section 3 stands: CHWP-03 tripped on motor overload at 02:14 on 4 March, third occurrence in nine weeks, 3 hours 40 minutes of lost cooling to Block C. The first two occurrences had been closed in the maintenance system as "motor overload, reset, returned to service". That closure text is the whole problem in miniature: it records the symptom and the response, and nothing that would help anybody a month later.

On the third occurrence, evidence was preserved. The drive's own fault log held current traces for all three trips and showed a slow rise in running current over the preceding weeks in each case, resetting after each restart. The bearing housing was retained. Thermal images taken before shutdown showed a hot outboard bearing. The lubrication task instruction, when it was actually read, specified "grease bearings" with no grease type and no quantity.

A 5 Whys on that evidence runs cleanly. Why did the motor trip on overload? Because the driven load increased as the outboard bearing degraded. Why did the bearing degrade? Because it ran with insufficient lubrication. Why was lubrication insufficient? Because the task specified neither quantity nor grease type, so the amount applied varied with whoever did the work. Why was the instruction written that way? Because the PM library had been migrated from a spreadsheet during a system implementation with no technical review of task content. Why no review? Because the migration scope treated task text as data to be transferred rather than content to be validated, and no engineering sign-off was required at cutover.

Note where that chain stops being about this pump. The last two answers are about the PM library and the migration, and that is where the leverage sits: the same failure mode is almost certainly latent across every asset class whose tasks came through that migration. Note also what the report would have said if the investigation had stopped at answer one: "motor overload due to increased load", action "replace motor and monitor". That is the version written when evidence is not preserved, and it would have delivered a fourth trip.

Example B: the standby generator that failed to take load. Hypothetical again: during monthly off-load tests the standby generator started normally, but during a live transfer it started and then shut down on low oil pressure after roughly forty seconds. The tempting cause is "oil pressure sensor faulty", and the sensor did read low. But tested against the evidence, the retained oil sample was below minimum level and visibly degraded, while the test record showed twelve consecutive satisfactory off-load runs.

None of the causal factors that emerged is a faulty part. The test regime only ever ran the machine off load, which does not exercise the lubrication system under load and so does not reveal a marginal condition. The oil level check was in the task instruction but recorded as a tick rather than a measured value, so a slow decline over a year was invisible in the record. And the condition was only discoverable during a real transfer: a classic hidden failure, where the protective function worked exactly as designed while the thing it protected against developed unobserved. The corrective actions are a test regime that loads the machine, a recorded quantitative oil level rather than a tick, and a review of which other protective functions are only ever tested in a way that would not reveal a developing fault. "Replace the sensor" would have closed the report and left all three conditions in place.

7. Human and organisational factors, without blame

The reason RCA so often stops at "operator error" is not analytical. Going further is uncomfortable, because the next level up is usually a decision somebody in the room made: a schedule that was too tight, a procedure nobody reviewed, a training budget deferred, a migration signed off without technical review. An investigation culture that treats naming those conditions as an accusation will never get past the individual, and will keep producing reports that identify the last person to touch the equipment.

It helps to be able to point at the discipline's own scope here. IEC 62740:2015 explicitly excludes the assignment of responsibility or liability from RCA. That is a useful thing to say at the start of an investigation, because it separates two activities organisations habitually run together and should not. Determining accountability, where that is necessary, is a management or legal process with its own rules and protections. RCA is a technical process aimed at preventing recurrence. Running them as one corrupts both: the analysis gets shaped by its consequences for individuals, and people stop telling you things.

The practical consequences are concrete rather than philosophical. Say in the opening meeting that the purpose is preventing recurrence and that accountability is a separate process. Write causal statements about conditions rather than people: not "the technician failed to check the oil level" but "the task recorded oil level as a tick rather than a measured value, so a gradual decline was not visible in the record". Where an action was genuinely a deviation, ask what made the deviation attractive or necessary, because if it saved time or avoided an obstruction then anyone in that position will do the same again. And separate what happened from who is accountable in time as well as in principle, so the technical analysis is complete before any accountability discussion begins.

One jurisdictional note for readers in regulated process environments. In the United States, the process safety management regulation at 29 CFR 1910.119 requires incident investigation for covered processes, with its own requirements on timing and documentation. That is a US federal requirement only and has no force in the UAE, the wider Gulf or the UK, where the governing instruments are entirely different. Where an investigation may feed a regulatory process, involve whoever handles that early, because the sequencing and documentation requirements will shape what you can do with your findings.

8. Telling a real corrective action from a restatement

The last third of an RCA report is where good analysis most often dies. The causes are sound and then the actions are "reinforce the importance of following procedure", "remind staff to check oil levels", "increase awareness", "monitor closely". None changes a condition. They are the problem statement rewritten in the imperative mood, and they cannot fail an implementation check because there is nothing to check. The table below contrasts weak and strong statements at both stages, causal and corrective, and it is the single most useful page to put in front of a team learning to write these reports.

Weak Why it fails Strong
Cause: the bearing failed. Restates the event. No condition named, nothing to remove. Cause: the outboard bearing ran under-lubricated because the PM task specified no grease quantity or type.
Cause: operator error. Describes an action and offers it as an explanation. Ends the investigation at the useful point. Cause: two identical adjacent valves are unlabelled and operate in opposite senses, so selection depends on memory.
Cause: lack of maintenance. True of nearly every failure and therefore says nothing. Not testable. Cause: the PM was deferred four times in six months because planned duration was 20 minutes against an actual 40.
Cause: inadequate training. Names a deficiency without naming the task, the gap or who assessed it. Cause: no task-specific competency assessment exists for setting this drive's overload parameters; induction records were treated as sufficient.
Action: remind staff to follow the procedure. Changes no condition. Cannot be verified or measured. Unfalsifiable. Action: revise task PM-0412 to state grease type and quantity per bearing; reissue and confirm the old revision is withdrawn. Owner, date.
Action: monitor the equipment closely. No parameter, no threshold, no responder, no end date. Action: trend motor running current weekly against a defined alarm threshold, with a named responder and a defined response. Owner, date.
Action: increase awareness of the issue. Awareness is not a control. Decays within weeks and leaves no trace. Action: label both valves to the site labelling standard and fit a physical interlock preventing the incorrect sequence. Owner, date.
Action: replace the failed component. Necessary repair, not a corrective action. Restores service and leaves the cause in place. Action: replace the component (repair), and separately review the 34 other tasks migrated in the same batch for missing technical content. Owner, date.

Three habits make the difference. Keep repair separate from corrective action in the report, because merging them lets the repair stand in for prevention. Require every action to name a condition it changes, and reject any that cannot. And prefer actions higher up the control hierarchy where the consequence justifies it: eliminating or engineering out a condition outlasts a procedural or training control, which depends on sustained attention to keep working. Procedural controls are not worthless, and sometimes they are all that is available, but they should be a conscious choice rather than the default. The formal expression of that preference is the hierarchy of controls, as required by ISO 45001:2018 clause 8.1.2 and described by NIOSH. For the management-system side of tracking, verifying and closing these actions, the CAPA guide on corrective and preventive action covers the machinery.

9. How deep to go, and when to stop

"Root" is a misleading word. It suggests a single terminal cause waiting to be found, which is rarely the shape of reality. Most failures have several contributing conditions at several depths, and the chain does not terminate so much as gradually leave your sphere of influence. The real skill is not finding the bottom, it is choosing the right level to act at. The heuristics I would use, in order of usefulness:

  • Stop at the deepest cause you can actually control. If the next answer up the chain is about market conditions or national skills supply, you have gone past the useful level. Come back one step.
  • Match depth to consequence. A failed corridor light and a failed life-safety system warrant very different investigation depth. Spending a week on the former is not thoroughness, it is misallocation, and it is why organisations with a blanket "RCA on everything" rule end up doing RCA on nothing properly.
  • Ask whether the cause is systemic. If a cause plausibly affects other assets, other sites or other tasks, that is the level to act at, because the leverage is there. Example A's real value was not the pump, it was the migrated task library.
  • Stop when further depth produces no new action. If two more why-steps lead to the same corrective action as the previous one, they are commentary rather than analysis.
  • Accept multiple causes. If three conditions all had to be present, say so and act on all three. Forcing a single root cause because the template has one box is a documentation failure masquerading as rigour.
What RCA costs, and where it does not pay

A proper investigation on a significant failure is days of competent people's time, not hours, plus laboratory work where the mechanism is physical. That cost is easy to justify on a high-consequence or recurring failure and impossible to justify on a low-consequence one-off. Organisations that mandate full RCA on every failure do not get more rigour; they get a form-filling culture that produces shallow reports on everything, including the failures that deserved real work. Set a trigger threshold on consequence and recurrence, be explicit that failures below it get a repair and a well-coded history record rather than an investigation, and defend the difference.

10. How RCA programmes fail

The patterns are consistent enough across organisations to be worth naming plainly. Every one is a management and discipline problem rather than a technical one, which is encouraging: the fix is in the organisation's own hands.

  • RCA run to satisfy a form. A template exists, the boxes get filled, the report is filed, nothing changes. The tell is a completed five-why grid whose answers are all restatements and an action section full of reminders. The cause is a mandate to complete RCAs combined with no review of their quality.
  • Stopping at the first plausible cause. Someone experienced offers a mechanism in the first hour, it sounds right, and the rest of the investigation becomes a search for supporting detail. The most damaging pattern, because the output looks entirely professional. The defence is step 5.
  • Actions nobody owns. Actions assigned to a department, or to "maintenance", or with no date, do not happen. Assigned to a named person with a date they sometimes do not happen either, but at least the omission is visible.
  • Verifying implementation but never effectiveness. The register shows every action closed and the failure is still recurring. Effectiveness needs a measure and a review date set when the action is written, not decided later.
  • Evidence lost before anyone started. The investigation convenes a week later, the part is gone, the trend data has been aggregated, and everybody has agreed on a story. The output is a guess with a report cover.
  • Blame leaking in. Once one investigation has led to a disciplinary outcome, the quality of witness evidence in every subsequent investigation falls permanently. Slow damage, very hard to reverse.
  • Findings that never leave the report. A cause identified as systemic, acted on for the one asset investigated, never applied across the others it affects. Pure waste: the expensive analytical work is done and the cheap part, generalising it, is skipped.

The idea to walk away with

Root cause analysis is not a diagram and it is not a template. It is the discipline of refusing to stop at a cause you cannot act on. Everything else follows. The precise problem statement exists so you know what you are explaining. The evidence discipline exists so candidate causes can be tested rather than merely believed. The timeline exists to expose gaps and kill impossible theories. The methods exist to widen or structure the search. And the corrective actions exist to change a condition, which is the only thing that ever prevents a recurrence.

If you take one operational change from this guide, make it evidence: preserve the part, export the data at full resolution, photograph before anything is touched, interview individually and early. An investigation with good evidence and a crude method beats an elegant analysis built on recollection every time. If you take one writing change, make it the causal statement test.

Final thoughts

The gap between organisations that learn from failures and organisations that keep having the same ones is not analytical sophistication. Both groups know what a fishbone diagram is. The difference is that one group preserves evidence before restoring service, tests its favourite theory instead of confirming it, writes actions that change conditions, checks months later whether the failure actually stopped, and applies what it learned to the other assets with the same exposure. None of that requires a licence, a certification or a platform. It requires that somebody senior enough reads the reports and sends back the ones ending in "operator error" or "reminded staff".

Start there. Take the last ten closed RCA reports in your organisation, apply the causal statement test to each conclusion and the condition test to each corrective action, and count how many survive. That number tells you more about your reliability trajectory than any KPI dashboard, and improving it costs nothing but attention.

Disclosure

Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.

Reviewing how your organisation investigates failures?

Independent advisory on investigation process design, failure coding and history capture, corrective action tracking, and getting reliability data into a state where recurrence questions are actually answerable. 22+ years across utilities, oil and gas, manufacturing, government and facility operations.

Book a conversation

Related reading: The 5 Whys method, Fishbone diagrams, RCA tools: 8 methods explained, RCA examples from equipment failures, Fault tree analysis, Pareto analysis, 8D problem solving, CAPA explained, Incident investigation, Failure analysis methods, Failure codes: Problem, Cause, Action.

Muhammad Abbas

CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.

Work with me
MAbbaz.com
© MAbbaz.com