mail@mabbaz.com Abu Dhabi, UAE

Root Cause Analysis · Reliability · Tool Selection

Root Cause Analysis Tools: 8 Methods Explained

Most root cause analysis goes wrong before anybody draws a diagram, because the wrong tool was picked for the shape of the problem. This is a selection guide: eight recognised RCA methods, what each one is good at, where each one fails, and a framework for choosing and combining them.

Muhammad Abbas September 27, 2026 ~19 min read

There is a particular kind of meeting that plays out far too often. A pump has failed for the third time in a year, somebody has been asked to do a root cause analysis, and the output is a 5 Whys form with five boxes filled in and a final answer of "inadequate maintenance". The form is complete. The investigation is worthless. Not because 5 Whys is a bad method, but because it was the wrong method for a problem with several contributing causes running in parallel, and it was the one chosen because it was the template attached to the corrective action procedure. Tool selection is the quiet, unglamorous decision that determines whether an RCA produces an insight or a tick in a box.

The message up front: there is no best RCA tool, only a best fit for the shape of the problem in front of you. Single causal chain, one event, technical failure: something linear and fast. Multiple contributing causes, several actors, time-dependent conditions, protection layers that should have caught it: something structured and visual, usually more than one method in sequence. Choose by problem shape, not by what the template mandates.

1. Why the tool choice decides the outcome

Every RCA method is a way of organising evidence and reasoning. None of them generates evidence. None of them supplies judgement. What each one does is impose a particular shape on the analysis, and that shape is either congruent with the problem or it distorts it.

A linear method forces you to pick one cause at each step and follow it. If the failure really was a single chain, that is efficiency. If the failure had four contributing causes that only mattered in combination, the linear method silently discards three of them, and the discarding happens invisibly, inside somebody's choice of which answer to write in the next box. A tree method forces you to be explicit about whether causes combine or substitute, which is far more honest, but it costs a great deal more effort and it demands data you may not have. A timeline method reconstructs sequence beautifully and says almost nothing about causation on its own. A prioritisation method tells you where to look and never tells you why.

So the practitioner's skill is not mastery of eight techniques. It is recognising, in the first hour, what shape the problem has, and reaching for the method whose shape matches. That recognition is what this article is for. If you want the step-by-step process that sits around whichever tool you choose, the evidence gathering, the causal statement discipline, the corrective action set and the verification, that lives in the root cause analysis process pillar. This piece is purely about choosing.

2. What the standards actually say about RCA techniques

A surprising number of articles on this subject open by asserting that root cause analysis has no standard. That is not true, and it is worth correcting because the standard is useful.

IEC 62740:2015 "Root cause analysis (RCA)" is an international standard dedicated to the subject, adopted in Europe as EN 62740:2015. It sets out principles and process steps, and it describes a set of named techniques, including the "Why" method that most people know as 5 Whys, the Fishbone or Ishikawa diagram, events and causal factors charting, causes tree analysis and fault tree analysis. Two boundaries in its scope matter for how you use it. First, it addresses analysis after an event has occurred, which is why prospective methods such as FMEA sit outside it. Second, it explicitly excludes the assignment of responsibility or liability. That second point is the single most useful sentence a facilities or reliability manager can quote when an investigation starts drifting towards finding somebody to blame.

Two neighbouring standards fill in the rest of the landscape. IEC 31010:2019 "Risk management - Risk assessment techniques" is the catalogue of assessment techniques, and note the designation carefully: it is IEC 31010, not ISO 31010, and not ISO/IEC 31010. Several of the methods below appear in it as risk assessment techniques rather than as RCA techniques, which is itself a clue about what they are for. IEC 61025:2006 remains the current edition for fault tree analysis, and IEC 60812:2018, Edition 3, covers failure modes and effects analysis, with the title changed at that edition to name FMEA and FMECA directly.

What a standard does not give you

IEC 62740 describes techniques. It does not prescribe how to run a 5 Whys session, and no standard I am aware of does. So the correct framing for the light methods is this: no standard prescribes how to run it, but IEC 62740:2015 describes it as one of several recognised RCA techniques. That is a stronger position than either "it is standardised" or "it is just a whiteboard exercise". These documents are also paywalled and voluntary. They bind you through contract and internal procedure, not through law.

3. The eight methods at a glance

This is the table to keep. Read it as a routing table: find the row whose best-suited problem shape matches what you are looking at, then follow the link to the method's own guide for the mechanics.

Method What it is Best-suited problem shape Main weakness Read more
5 Whys Iterative questioning down a single causal chain One event, one obvious chain, low consequence, quick turnaround Cannot hold branches; forces a single answer at each step 5 Whys guide
Fishbone / Ishikawa Categorised cause map on a single diagram Several candidate causes, group session, need to open the space Generates hypotheses, not conclusions; no logic or weighting Fishbone guide
Fault tree analysis Top-down deductive tree with logic gates Combinations, redundancy, high-consequence systems Heavy effort; quantification needs data most sites lack FTA guide
FMEA / FMECA Bottom-up, prospective failure mode and effect analysis Design and strategy work before failure, not after Not an after-the-event RCA tool at all (see below) FMEA guide
Pareto analysis Frequency or cost ranking to concentrate attention Recurring patterns across a population; choosing a target Prioritisation only; says where to look, never why Pareto guide
ECF charting / timeline Chronological reconstruction with conditions attached Incidents with many actors and time-dependent conditions Establishes sequence, not causation; labour intensive Incident investigation guide
Change analysis Systematic comparison of now against when it worked "It used to be fine" problems with a datable turning point Needs a genuine known-good baseline to compare against Worked RCA examples
Barrier analysis Audit of the defences that existed, failed or were absent Anywhere protection layers, permits or interlocks matter Tells you what failed to stop it, not what started it Failure analysis methods

4. The two light methods: 5 Whys and Fishbone

5 Whys is a linear interrogation. You state the problem, ask why it happened, accept an answer, and ask why of that answer, repeating until you reach something worth acting on. Its virtue is speed and accessibility: a supervisor can run it at the asset, in twenty minutes, with no training and no software. For a genuinely simple problem with one dominant chain, that is exactly proportionate, and the fact that it needs no specialist is what makes it the only method some organisations ever actually use.

Its weakness is structural and it is the most important single limitation in this article. The method has no way to hold a branch. At every step it demands one answer. When a failure had three contributing causes, the person filling in the form picks whichever one is most available to them, and the other two disappear without ever being recorded as rejected. You do not get a wrong answer so much as an arbitrary one. Reach for 5 Whys when you are confident the chain is single and you want speed. Reach for something else the moment somebody in the room says "well, it could also have been".

Fishbone, or Ishikawa, solves precisely that problem, and only that problem. It is a categorised map that holds many candidate causes in view at once, grouped under headings such as people, method, machine, material, measurement and environment. In a group session it is unmatched at surfacing what a team collectively knows and at preventing the loudest person from closing the analysis early, because every suggestion gets a bone rather than a debate.

What it does not do is conclude. A completed fishbone is a well-organised list of hypotheses with no logic, no weighting and no evidence attached. Treating one as a finding is a common and costly mistake, because the diagram looks like an analysis. It is a hypothesis generator, which makes it an excellent second step and a poor last one. Use it to open the candidate space, then test the candidates with evidence and drill the surviving ones with something linear.

5. Fault tree analysis: when combinations matter

Fault tree analysis inverts the direction of travel. Instead of working forwards or sideways from symptoms, you state the undesired top event and decompose it downwards through logic gates into the combinations of lower-level events that could produce it. The AND and OR gates are the whole point: they make explicit whether two conditions had to coincide or whether either alone was sufficient. No other method on this list expresses that distinction cleanly, and that is precisely the distinction that matters in any system with redundancy, standby equipment, voting logic or layered protection.

If you have a duty and standby pump set and the plant still lost supply, a fishbone will list plausible causes and a 5 Whys will pick one. A fault tree will make you say out loud that the outcome required the duty pump to fail AND the standby to fail to start, and will then force the question of whether those two failures shared a cause. Common cause failure is where redundancy quietly stops being redundancy, and fault trees are how you find it. IEC 61025:2006 remains the current edition governing the technique.

The cost is real. Building a defensible tree takes engineering time, system knowledge and discipline about scope, and the quantitative version needs failure rate data that most facilities and utilities simply do not hold at the required quality. My advice is to use fault trees qualitatively far more often than quantitatively. The structure of the tree, the logic of what had to combine, delivers most of the insight; the probability arithmetic is the part that stalls. Reserve the method for consequences that justify the effort, and accept that a tree built on invented failure rates is worse than no tree, because it is confidently numeric.

6. FMEA and FMECA: the distinction that is constantly missed

FMEA appears on almost every published list of root cause analysis tools, and on most of them it does not belong. This is worth stating plainly because the confusion causes real wasted effort.

FMEA works bottom up and, crucially, it works prospectively. You take a component or a function, enumerate the ways it could fail, trace the effects of each failure mode upwards, and rank them so that you can act before anything has happened. FMECA adds a criticality dimension to that ranking. The governing international standard is IEC 60812:2018, Edition 3. Strictly, both are risk analysis methods, not after-the-event RCA tools. IEC 62740 confines itself to analysis after an event has occurred, which is exactly why the prospective methods are not among its techniques.

The test that separates them

Ask what has already happened. RCA starts from a real event and reasons backwards to why it occurred. FMEA starts from a system that has not failed yet and reasons forwards to what could occur. If you are holding a broken pump, an FMEA will not tell you why it broke. If you are deciding which pumps deserve which maintenance strategy, an FMEA is the right instrument and an RCA has nothing to offer.

The two do connect, and the connection is where the value sits. A completed RCA that identifies a failure mode nobody had anticipated should feed back into the FMEA, which in turn should reshape the maintenance strategy. Run in that loop, FMEA is the mechanism by which one investigation improves the whole programme rather than just fixing one asset. For the mechanics of the method see the FMEA guide, and for how failure mode thinking links to remaining life modelling, from RUL to FMEA.

7. Pareto analysis: prioritisation, not causation

Pareto analysis ranks a population of events by frequency, cost or downtime and shows you that a small number of categories account for most of the total. On a maintenance dataset it typically shows a handful of asset classes or failure codes generating the bulk of unplanned work. It is cheap, it uses data you already own, and it is one of the more reliable ways of deciding what to investigate when everything is nominally a priority.

It is not, however, a causal method, and it is routinely mistaken for one. A Pareto chart showing that bearing failures dominate your pump work orders does not tell you why bearings fail. It tells you that bearings are where the causal question is worth asking. That is enormously valuable and completely different. Pareto is the front door to an RCA, not the RCA.

Two practical cautions. First, the output depends entirely on how the categories were coded, so a Pareto built on inconsistent failure coding will faithfully rank your data entry habits rather than your failures. Getting the coding structure right is the precondition, which is why I treat failure code structure as part of the RCA toolkit rather than a separate administrative topic. Second, ranking by count and ranking by consequence give different answers, and the consequence ranking is usually the one that should drive investigation effort. Run both. Where they disagree, that disagreement is itself informative. The method's full mechanics, including the cumulative curve and the category traps, are in the Pareto analysis guide.

8. Timeline methods and change analysis

Events and causal factors charting, and timeline analysis more generally, reconstructs what happened in order, with the conditions that applied at each point attached alongside. It is named among the techniques described in IEC 62740. Where it earns its place is in incidents with many actors and time-dependent conditions: a handover that happened at the wrong moment, an alarm acknowledged by one person and acted on by another, a control isolated for work that was then reinstated out of sequence. Human accounts of such events are commonly reordered in memory. Building the chart against timestamped evidence, from logs, alarm histories, access records and work order timings, is what turns anecdote into a defensible sequence.

Its limitation is that sequence is not causation. A complete timeline still requires you to reason about which events and conditions were causal, which is why timeline work usually pairs with a causal method rather than standing alone. It is also the most labour intensive of the light methods, because most of the work is evidence assembly rather than analysis.

Change analysis is the most underused method on this list and often the highest yield per hour invested. The question is deceptively simple: it worked before and it does not work now, so what is different? You compare the failed situation against a known-good reference across every dimension you can enumerate, including the equipment itself, the spare parts and their suppliers, the lubricant, the operating regime and duty, the software or controller version, the setpoints, the environment, the people and their training, the procedure revision, and the maintenance interval.

The reason it works so well is that most sustained new problems in a stable system were introduced by a change, and changes are far easier to enumerate than causes are to imagine. A pump that ran for eight years and started failing every four months did not develop a new physics; something changed. Change analysis regularly finds in an afternoon what two rounds of fishbone workshops missed, because the fishbone was asking the team to imagine causes while change analysis asked them to list facts. Its one requirement is a genuine known-good baseline. If the asset has never been reliable, there is nothing to compare against, and the method has nothing to offer.

9. Barrier analysis, and where bowtie sits

Barrier analysis asks a different question from every other method here. Rather than what caused the event, it asks what should have stopped it. You identify the hazard and the target, enumerate the barriers that were supposed to sit between them, and classify each one: it existed and held, it existed and failed, it existed but was bypassed, or it was never there at all. The barriers include physical guards and interlocks, procedural controls such as permits and isolations, and administrative ones such as competence checks and supervision.

This is the method to reach for wherever protection layers matter, which in practice means most safety-relevant incidents and a good deal of high-consequence equipment damage. It is unusually good at producing corrective actions that are genuinely preventive, because a missing or bypassed barrier is a concrete, buildable thing, whereas a cause expressed as "inadequate awareness" is not. It also has a useful side effect: it reframes the analysis away from who made the error and towards why the system permitted the error to reach the target, which is the posture IEC 62740 expects when it excludes the assignment of responsibility.

Its limitation is the mirror image of its strength. Barrier analysis tells you what failed to stop the event and not what started it. Used alone it can produce an action set that adds a protection layer while leaving the initiating cause entirely untouched, which is how organisations accumulate barriers and keep having the same event.

Worth naming alongside it: bowtie mapping, which places a hazardous top event in the centre with threats and preventive barriers on the left and consequences with mitigating barriers on the right. It is closely related in thinking and genuinely useful for laying out a protection scheme. Two accuracy points. Bowtie is forward-looking risk mapping rather than after-the-event analysis, and it is a 2018 CCPS and Energy Institute concept book, not a standard. There is no bowtie standard, whatever a training brochure says, although bowtie does appear as a listed technique in IEC 31010:2019.

10. 8D and the difference between a tool and a container

8D belongs in this discussion, but not in the same category as the rest. It is a structured problem-solving discipline in eight ordered steps, running from team formation and interim containment through root cause identification, corrective action, verification and prevention, to closure. What it does not supply is a method of causal reasoning. The root cause step in an 8D is performed using one or more of the tools above.

That makes 8D a container process rather than a tool, and understanding the distinction stops a specific argument that wastes time. "Should we use 8D or 5 Whys?" is a malformed question. 8D is the wrapper that ensures containment happens before analysis, that the analysis is verified, and that the fix is proven and institutionalised rather than assumed. 5 Whys is one of the things you might do inside step four. If you work into automotive or manufacturing supply chains you will be asked for 8D by format, and the right response is to choose the causal method that fits the problem and run it inside the 8D structure. See the 8D problem solving guide for the eight steps in full.

11. The selection framework

Here is the useful part. Rather than memorising eight methods, ask six questions about the problem and let the answers route you.

  • Single chain or multiple contributing causes? If the causal path is plausibly one line, a linear method is proportionate. The moment you suspect causes that only mattered in combination, you need something that can hold branches: fishbone to open the space, fault tree if the combinations are the crux.
  • Technical or organisational? Purely technical failures respond well to fault tree and to failure analysis of the physical evidence. Failures with procedural, competence or communication content need methods that can carry those categories, which means fishbone, barrier analysis and timeline work. A method that has no place to put "the procedure revision was never issued to the night shift" will quietly drop it.
  • One event or a recurring pattern? A single event is an investigation. A recurring pattern is first a data question, so start with Pareto or trend analysis across the population, then investigate the concentration you find. Running a full RCA on one instance of a recurring problem, without looking at the population, is how you fix the instance and keep the pattern.
  • Does timing matter? If the order or coincidence of events is part of the story, if things happened during a handover, a startup, a shutdown or an alarm flood, build the timeline before you do anything else. Without it you will reason about a sequence you have partly imagined.
  • Were protection layers involved? If a permit, interlock, guard, alarm, relief device or standby system should have prevented or limited the outcome, barrier analysis is not optional. Its absence is the usual reason an action set looks thin.
  • How much consequence justifies how much effort? This is the discipline question. A quantified fault tree on a failed corridor light is waste. A 5 Whys on a fire pump that failed to start on demand is negligence. Match analytical depth to the consequence of the event and to the consequence of getting the answer wrong.

Answer those six and the method usually picks itself. Where the answers pull in different directions, that is the signal to combine methods rather than to compromise on one.

12. How the tools combine in sequence

Experienced investigators rarely use one method. They use a short sequence in which each step feeds the next: prioritise, establish, open, drill, test, act. The full sequence looks like this.

Pareto to choose the target. Rank the population by frequency and by consequence and pick the concentration worth the effort. Timeline to establish what happened. Build the sequence from timestamped evidence before anybody theorises, because theories contaminate recollection. Fishbone to open the candidate space. Get every plausible cause on one page, including the ones you will discard, so that the discarding is visible. Change analysis in parallel if a known-good baseline exists, because it often shortcuts the whole exercise. 5 Whys to drill the surviving branch once evidence has narrowed the candidates to one or two. Barrier analysis to test the defences, asking what should have stopped this and why it did not. A documented action set at the end, each action tied to a specific identified cause or missing barrier, with an owner, a date and a verification step. Bring in a fault tree in place of the fishbone and 5 Whys where combinations and redundancy are the heart of the problem and the consequence justifies the effort.

Problem shape Recommended sequence Why this order
Simple single-chain failure, low consequence 5 Whys, then documented actions Proportionate. Anything heavier is waste and will not get done.
Repeat failure of one asset class across a site Pareto, then change analysis, then fishbone, then 5 Whys on the surviving branch Population first, then look for an introduced change before imagining causes.
"It used to be reliable" with a datable turning point Change analysis first, then 5 Whys to confirm the mechanism Enumerating differences is faster and more reliable than imagining causes.
Safety incident with multiple people involved Timeline, then barrier analysis, then fishbone for organisational factors Sequence before causation; barriers produce preventive actions.
Redundant or protected system that failed anyway Fault tree, qualitative first, with common cause specifically examined Only a tree expresses whether failures had to combine or shared a cause.
Time-dependent event during handover, startup or alarm flood Timeline, then fishbone, then barrier analysis Timing is the substance of the problem, so reconstruct it first.
Customer or contractual complaint requiring a formal response 8D as the container, with the causal method chosen by shape inside step four The format is mandated; the reasoning method still has to fit the problem.
Deciding strategy for assets that have not failed yet FMEA or FMECA, not RCA at all Prospective question, prospective method. RCA has nothing to say here.

These are illustrative routings drawn from the kind of problems I meet in facilities, utilities and plant maintenance work, not prescriptions. Adapt them to your own event types. For fully worked examples of these sequences applied to equipment failures, see the worked RCA examples, and for the physical evidence side of a technical investigation, failure analysis methods.

13. Three ways tool selection goes wrong

Using the tool the template mandates rather than the one the problem needs. This is the most common failure by a wide margin. An organisation adopts a corrective action procedure with a 5 Whys form attached, and from that day every problem is a single-chain problem because the form only has one chain. The fix is not to abandon the template but to make the method a deliberate choice recorded at the start of the investigation, with a line that says which method was selected and why. That one line changes behaviour, because it makes an unconsidered default visible.

Using a heavyweight method as a substitute for evidence. A fault tree built from assumption rather than measurement is not more rigorous than a 5 Whys, it is less, because its apparent rigour discourages challenge. Trees with quantified gate probabilities whose underlying failure rates were guessed in a meeting are common. The formality of the method does not upgrade the quality of the inputs. If the evidence is thin, the honest output is a shortlist of hypotheses and a plan for the tests that would distinguish them, not an elaborate diagram.

Using any of them without evidence discipline. Every method on this list will happily process opinion. None of them contains a mechanism that rejects an unsupported claim. The discipline has to come from the investigator: each causal statement gets an evidence reference, or it is explicitly marked as unverified with a test that would settle it. An RCA where nobody can say how a claim was established is not an analysis, whichever diagram it produced.

What none of these tools can fix

Tool choice cannot rescue an investigation that lacks time, access to evidence, or the authority to act on the finding. If the analysis is expected in an hour, if the failed part has already been scrapped, or if the likely cause sits with a decision nobody in the room can revisit, no method will produce a useful outcome. Those are organisational conditions, not analytical ones, and they are worth naming before the investigation starts rather than explaining afterwards. Software helps with the record keeping and the trending, and nothing more; a workflow in a maintenance system does not supply reasoning.

The idea to walk away with

The eight methods are not competing products and they are not ranked by sophistication. They are instruments with different shapes, and the shape of the problem decides which one fits. Single chain: something linear. Multiple candidates: something visual. Combinations and redundancy: something deductive. A pattern rather than an event: something statistical first. Time-dependent: a timeline. Protection layers: barriers. A baseline that used to work: change analysis. Not failed yet: not RCA at all.

And the strongest investigations chain them. Pareto to choose, timeline to establish, fishbone to open, 5 Whys to drill, barrier analysis to test, documented actions to close. That sequence costs little more than the single-method habit and produces findings that survive a review.

Final thoughts

If I could change one thing about how root cause analysis is commonly practised, it would not be the adoption of a more sophisticated method. It would be the insertion of a single explicit decision at the start: what shape is this problem, and therefore which method am I going to use. Most organisations have the tools already, in training decks nobody opens, and use one of them for everything.

The two additions I would make to a standard toolkit are change analysis and barrier analysis, because they are cheap, they need no specialist training, and between them they cover the two questions the popular methods tend to skip: what is different now, and what should have stopped this. Add those, make the method choice deliberate, hold the line on evidence, and the quality of your investigations will improve more than any new template could deliver. The relevant reference if you want to read the recognised techniques in their authoritative form is IEC 62740:2015, available through the IEC , and the bowtie concept book through the CCPS at AIChE . Both are paid publications, and both are voluntary references rather than law in any jurisdiction.

Disclosure

Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.

Choosing the right RCA method for a recurring failure?

Independent advisory on RCA method selection, failure coding structure, investigation governance and the corrective action loop that makes findings stick. 22+ years across utilities, oil and gas, manufacturing, government and facility operations.

Book a conversation

Related reading: Root cause analysis: methods and step-by-step guide, 5 Whys, Fishbone diagrams, Fault tree analysis, Pareto analysis, 8D problem solving, Incident investigation, Failure codes: Problem, Cause, Action.

Muhammad Abbas

CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.

Work with me
MAbbaz.com
© MAbbaz.com