mail@mabbaz.com Abu Dhabi, UAE

Root Cause Analysis · Reliability · Method Guide

5 Whys Root Cause Analysis: Method and Examples

The 5 Whys is the most widely used root cause analysis method in the world and the most widely misused. It is quick, it needs no software, and in the right hands it gets past a symptom to something you can actually fix. In the wrong hands it produces a tidy linear story that is wrong, and often ends by blaming a technician. This is the method taught properly, with worked chains, a liftable template, and an honest account of where it breaks.

Muhammad Abbas September 27, 2026 ~18 min read

Ask a maintenance team how they do root cause analysis and the answer, nine times in ten, is "we do a 5 Whys". Ask to see the last one and you will usually be shown a form with five lines on it, the fifth line reading something like "operator error" or "the technician did not follow the procedure", and a corrective action that says "retrain". That is not a 5 Whys. That is a blame form with a 5 Whys layout. The method itself is genuinely good, genuinely cheap, and genuinely capable of finding fixable causes, but only if you run it with two disciplines most people skip: each why must interrogate the previous answer rather than the original problem, and every link in the chain must be supported by evidence rather than by the loudest opinion in the room.

The message up front: the 5 Whys is a single-chain method. It is excellent for a simple, well-bounded failure where one causal path dominates, and it is the wrong tool the moment a failure has several contributing causes running in parallel, because the linear format forces you to pick one branch and silently discard the others. Learn to recognise that moment and escalate to a fishbone or a fault tree. Knowing when to stop using the 5 Whys is as much a part of the skill as knowing how to run one.

1. Where the 5 Whys came from, and what it was built for

The method's lineage runs through the Toyota Production System. It is credited to Sakichi Toyoda and was developed further inside Toyota's manufacturing practice, where Taiichi Ohno described repeated questioning as the way to get past a symptom to the underlying condition. That origin matters, because it tells you what the tool was designed to do and, by omission, what it was not.

It was designed for the shop floor. It was designed to be run at the place the problem happened, by the people who saw it happen, within minutes of it happening, without a facilitator, a licence or a wall of sticky notes. It was designed to prevent the standard reflex of fixing what you can see and walking away. In a production system built on standardised work, most defects have a reasonably traceable single path back to a process condition, and for that class of problem a short chain of whys is remarkably effective.

What it was not designed for is a complex multi-factor failure in a system with redundancy, human factors, environmental variability and latent design weaknesses all interacting. That is not a criticism of Toyota. It is a criticism of everyone who took a lean shop-floor tool, put it on a corporate incident form, and made it the mandatory answer for every failure regardless of complexity.

On the question of standing: no standard prescribes how to run a 5 Whys, but it is not a folk method with no formal recognition either. IEC 62740:2015 "Root cause analysis (RCA)" describes the "Why" method, which is what most of us call the 5 Whys, as one of several recognised RCA techniques, alongside the Fishbone or Ishikawa diagram, cause tree and fault tree approaches among others. The standard is published by the IEC and is paywalled, so I will not quote its wording. Two things about it are worth carrying into practice. First, it treats RCA as after-the-event analysis with a defined process around it, not as a single tool. Second, and usefully for the argument later in this guide, it explicitly excludes assigning responsibility or liability from the scope of root cause analysis. If you have ever needed an authority to point at when someone tries to turn an investigation into a disciplinary exercise, that is it.

For the wider landscape of methods and when each one applies, the companion piece is the root cause analysis methods and step-by-step guide. This guide deliberately stays inside the 5 Whys.

2. How to run one properly

There are four mechanics that separate a real 5 Whys from a filled-in form. None of them is difficult and all of them are routinely skipped.

Start from a precisely stated problem. The quality of the chain is capped by the quality of the opening statement. "Pump failure" is not a problem statement, it is a category. "Chilled water pump CHWP-03 tripped on high motor current at 09:40 on the third consecutive Monday, leaving the east wing without cooling for four hours" is a problem statement. It contains the object, the observed condition, the timing pattern and the consequence. Every one of those details gives the chain something to bite on. A vague statement produces a vague chain, and a vague chain produces a corrective action like "improve maintenance".

Ask why about the previous answer, not the original problem. This is one of the most common structural errors and it is worth labouring. If the problem is "the pump tripped" and your five whys read "why did it trip? bearing failure. why did it trip? no lubrication. why did it trip? the PM was missed", you have not built a chain, you have built a list of five parallel possible causes stacked vertically. A chain means why two interrogates the answer to why one. Each answer must be the direct, sufficient cause of the statement immediately above it, and you should be able to read the finished chain upwards as a sentence: because A, therefore B, therefore C. If it does not read upwards, it is not a chain.

Stop when you reach a cause you can actually control. The chain terminates not at a fixed depth but at the point where you reach a condition inside your organisation's control that, if changed, would prevent recurrence. Keep asking past that point and you walk out through the factory gate into causes you cannot act on: the supplier's quality system, the market, the weather, the founding budget decision. Those may be true and they are not useful outputs. A useful test: could a named person in this organisation own an action against this statement and change it within a reasonable horizon? If yes, you have arrived. If no, you have gone one why too far or one why too shallow, and you need to look at whether the previous level was actionable.

Ask the whys where the work happens. A 5 Whys run in a meeting room from a work order description produces the causes people can remember. One run at the asset with the technician who did the repair, the part in your hand and the trend on a screen produces the causes that are actually there. This is the part of the Toyota lineage that transfers best and gets dropped most often.

The reading-upwards test

Before you accept a completed chain, read it from the bottom to the top out loud, inserting "therefore" between each step. If any step does not follow from the one below it, you have either a gap in the logic or a jump across a branch. A chain that survives being read upwards is usually sound. Most chains do not survive it, which is exactly why the test is worth thirty seconds.

3. "Five" is a heuristic, not a rule

The number five is a memorable prompt to keep going past the first plausible answer. It is not a target and it is certainly not a quota. Treating it as one causes two distinct failures that I see in roughly equal measure.

Stopping at five when the chain is not finished. Some causal paths are genuinely longer. If the fifth answer is still a symptom that nobody can own an action against, keep asking. Six or seven whys is fine. The form has five boxes because someone designed a form, not because the causality obeys it.

Padding to five when the chain finished at three. This is the more insidious one, because it looks like diligence. A simple failure sometimes resolves properly in three steps. If you have reached a controllable cause at why three and the form demands five, people invent two more levels, and invented levels drift upward into abstraction: "why? because the culture does not prioritise maintenance". You have now replaced an actionable finding with an unactionable one, purely to satisfy a layout. Stop at three, write "chain complete at level 3" and move on.

The honest reframing: the method is "keep asking why until you reach something you can change, and be suspicious of yourself if that takes fewer than three steps". "5 Whys" is just the version of that sentence people remember.

4. The branching problem: the method's real weakness

This is the part that most 5 Whys training leaves out, and it is the most important thing in this guide. The 5 Whys produces a single linear chain. Real failures usually do not have a single linear cause.

Take a realistic hypothetical: a chilled water pump fails in service. The bearing was under-lubricated, and the greasing task had been deferred. But also, the vibration alarm that would have caught the developing fault had been disabled months earlier during commissioning and never re-enabled. And also, the standby pump did not auto-start because its control changeover had not been function-tested since installation. Three contributing causes, in three different domains: a maintenance execution gap, a monitoring configuration gap, and a redundancy verification gap. All three were necessary for the four-hour outage. Remove any one of them and the consequence is much smaller.

A single 5 Whys chain cannot hold that. The moment you write why one, you commit to a branch. Whichever cause the person filling the form happened to find first becomes "the" root cause, and the other two are not rejected on evidence, they are never written down at all. The corrective action addresses a third of the problem. Six months later the same outage happens through a different combination of the same three weaknesses, and the investigation concludes that the corrective action "was not effective".

There is a partial workaround: when an answer has more than one cause, split the chain and follow each branch separately. Done honestly this turns your 5 Whys into a small tree, at which point you should be honest that you have outgrown the tool and switch to something designed for trees.

When to stop using the 5 Whys and escalate

Switch tools when any of these is true: a single answer has two or more plausible causes that all look real; the failure crossed more than one discipline (mechanical plus controls plus operations); the consequence was safety-related, reportable or high cost; the same problem has already had a 5 Whys and recurred; or two competent people produce two different chains from the same facts. In those cases the honest move is to spread the causes out laterally in a fishbone diagram first, and if you need to reason about how causes combine logically, go to fault tree analysis, which is built for exactly the AND relationship that defeated the chain above.

Used in sequence these tools complement each other rather than compete. A fishbone is good at making sure no category of cause was ignored but weak at depth, because a fishbone bone stops at one or two levels. The 5 Whys is good at depth but blind to breadth. Running a fishbone to identify the candidate causes, then a short 5 Whys down each candidate that survives the evidence test, is a better default than either alone, and it costs maybe an extra hour. The wider set of options and where each fits is covered in root cause analysis tools: 8 methods explained.

5. The drift into blame, and how to spot it in one line

Here is a chain ending I have read many times in one form or another: "why did the seal fail? because it was installed incorrectly. why? because the technician did not follow the procedure." Full stop, action: retrain the technician, close.

That chain has stopped exactly one question early, and the missing question is the whole investigation. Why was the procedure not followed? And immediately behind it, was the procedure followable? Those two questions open up a completely different set of findings, and in practice they are usually where the real cause lives:

  • The procedure specified a torque value but the correct torque wrench was not in the van, and there was no route to get one without losing the shift.
  • The procedure was written for the superseded pump model and the current model's seal fits differently.
  • The procedure existed in a document library nobody can reach from a phone at the asset.
  • The procedure has eleven steps and the job is planned with thirty minutes of labour on it, so following it fully guarantees the technician is late for the next call.
  • Nobody has followed that step for three years, everyone knows it, and the deviation had become the de facto standard because it usually works.

Every one of those is an organisational condition that leadership can change. "The technician did not follow the procedure" is a description of the event, not a cause of it. It also happens to be the answer that terminates the investigation most comfortably for everyone senior in the room, which is precisely why it is so popular and why you should treat it as a red flag rather than a conclusion.

Two guardrails I would put in any RCA procedure. First, a rule that no chain may terminate on a named individual or on a statement of the form "person did not do X"; if it lands there, at least two more whys are mandatory. Second, the scope statement from IEC 62740:2015 quoted in spirit rather than in words: root cause analysis exists to find causes, not to assign responsibility or liability. Print that at the top of the form. It changes what people are willing to say in the room, and what people are willing to say in the room is the raw material of the whole method. This is the same logic that underpins a sound incident investigation process.

6. Evidence discipline: without it, a 5 Whys is a confident guess

A 5 Whys costs almost nothing to produce, and that is the danger. Five sentences written in a meeting room with no verification look identical on paper to five sentences each backed by a measurement, a photograph, a log entry or a witness statement. The document does not distinguish them. The reader cannot tell. Six months later nobody remembers which it was.

The fix is structural, not cultural: add an evidence column and refuse to accept a row without one. For each step, record what supports it and how strong that support is. Three honest categories are enough:

  • Verified: there is a physical artefact, measurement, log record, photograph or document that directly supports this statement, and it is referenced. The stripped thread was photographed. The BMS trend shows discharge pressure declining over nine weeks. The work order shows the PM was rescheduled twice and then cancelled.
  • Inferred: no direct evidence, but the statement is the only explanation consistent with what is verified, and you can say why the alternatives are excluded. Legitimate, but must be labelled.
  • Assumed: somebody in the room believed it. No supporting artefact, alternatives not excluded. Allowed to appear in a draft, never allowed to sit in a chain that a corrective action is built on.

The practical effect of this one column is out of all proportion to the effort. It stops the chain at the point where the knowledge actually runs out, instead of letting it run on confidently into fiction. And when a step comes back "assumed", that is not a failure of the analysis, it is the analysis telling you what to go and check. A 5 Whys with two verified steps and an honest "we do not yet know, here is how we will find out" is worth more than a complete five-step chain built on nothing.

This is also where the quality of your maintenance records stops being an administrative matter. A chain is only as verifiable as the history behind it, and if work orders are closed with a comment of "fixed" then most of your steps will be inference forever. Disciplined fault coding is what makes evidence available later, and the structure that does it is covered in failure codes: problem, cause, action.

7. Three worked chains, including one done badly

All three examples below are my own hypothetical equipment failures, constructed to illustrate the mechanics. The asset tags, dates and readings are invented. They are not drawn from any client engagement.

Example A (illustrative, hypothetical): air handling unit belt failure

Problem statement: AHU-12 supply fan stopped delivering air at 14:20 on 4 March. The drive belt was found sheared. The unit had been in service eleven months since the last belt replacement.

StepAnswerEvidence
Why 1The drive belt sheared in service.Verified: failed belt retained and photographed, clean shear with glazing on one face.
Why 2The belt had been running mistracked against the pulley flange, which abraded and then cut it.Verified: wear scar on pulley flange, matching abrasion pattern on belt edge.
Why 3The motor and fan pulleys were out of parallel alignment by a visible margin.Verified: straight-edge check at the asset before any disturbance, recorded with photograph.
Why 4Alignment was never checked after the previous belt change, because the PM task says "replace belt" and contains no alignment or tension step.Verified: PM task text in the CMMS, and the previous work order closure notes.
Why 5The PM task was written from the asset list at implementation without reference to the manufacturer's maintenance instruction, and has never been reviewed since.Verified: task creation date and unchanged revision history; OEM manual specifies alignment and tension check on reassembly.

Controllable cause reached: PM task content for the belt-drive AHU class is incomplete and unreviewed against OEM instructions. Action: revise the task to include alignment and tension verification with an acceptance criterion, and audit the rest of the belt-driven AHU population for the same omission. Note that the action has nothing to do with the technician who changed the belt. He did exactly what the task told him to do.

Example B (illustrative, hypothetical): the same failure done badly, then corrected

Here is a chain from the same hypothetical AHU, written the way these usually get written.

StepThe bad chainWhat is wrong with it
Why 1The AHU stopped working.This restates the problem instead of answering it. Level one is already wasted.
Why 2The belt broke.Answers the original problem, not the previous answer. Also no evidence recorded.
Why 3The belt was old.Assumed, and demonstrably false: eleven months on a belt is not age failure, and the shear pattern says otherwise. Nobody checked.
Why 4Maintenance was not done properly.Abstract, unevidenced, and a jump across an unexamined branch. It also aims the chain at people.
Why 5The technician did not do a thorough job.Terminates on an individual. Not a cause, and the investigation stops one question before the finding.

The resulting action, inevitably, is "toolbox talk on belt inspection" plus a note on someone's record. The asset fails again, because the PM task is still wrong and the next belt will be fitted mistracked by whoever is on shift.

The correction is Example A. Compare them step by step and the differences are mechanical rather than mysterious: the corrected chain opens with a specific problem statement instead of a category, each answer interrogates the one above it, every row carries evidence, no row names a person, and it terminates on a document the organisation owns and can change this week. Nothing about the corrected version required more skill. It required the discipline of the previous three sections.

Example C (illustrative, hypothetical): a chain that should have been abandoned

Problem statement: Standby generator GEN-02 failed to take load during the monthly black-start test on 2 June and the site remained on UPS until manual intervention at 09:12.

StepAnswerEvidence and note
Why 1The generator started but the changeover contactor did not transfer load.Verified: controller event log shows engine running, transfer command issued, no confirmation.
Why 2Branch detected. Either the contactor mechanism was seized, or the control signal never reached it, or the interlock inhibited transfer.Three candidates, all plausible, none excluded on evidence. This is the point to stop the chain.

At why two this analysis has branched, and the failure crosses mechanical, control and operational domains with a consequence that is close to safety-critical. Continuing as a single chain here would mean picking whichever branch the person running it finds easiest to argue and writing the other two out of history. The correct output of this 5 Whys is not a root cause. It is the decision to escalate: lay the three candidates out laterally, gather the evidence that excludes or confirms each one, and if they turn out to combine, model the logic in a fault tree. That is a successful 5 Whys. It cost ten minutes and it told you the method was insufficient before you built a corrective action on a guess. Examples of the same reasoning applied across other asset failures are collected in root cause analysis examples: equipment failures.

8. A 5 Whys template you can lift

This is the format I would put on a one-page form or a CMMS long-description field. The columns are deliberately awkward to fill in badly. Cause type is what stops the chain drifting: label each step as a physical condition, a human action, a system or process condition, or an organisational condition, and you will see immediately if your chain jumped from physical straight to human and stopped there.

Field Statement Evidence and status Cause type Action and owner
Problem statementAsset tag, observed condition, date and time, duration, consequence, and any pattern. One sentence, specific enough that a stranger could find the asset and the record.
Why 1Direct cause of the problem statement.What supports it. Verified / inferred / assumed.Physical conditionUsually none at this level.
Why 2Direct cause of Why 1.What supports it. Verified / inferred / assumed.Physical or processInterim containment, if any. Named owner.
Why 3Direct cause of Why 2.What supports it. Verified / inferred / assumed.Process conditionNamed owner, target date.
Why 4Direct cause of Why 3.What supports it. Verified / inferred / assumed.Process or organisationalNamed owner, target date.
Why 5Direct cause of Why 4. Extend to 6 or 7 if not yet controllable; stop earlier if it is.What supports it. Verified / inferred / assumed.Organisational conditionThe corrective action that prevents recurrence. Named owner, target date.
Branch checkDid any answer have more than one plausible cause? List them. If yes, record the decision to escalate to a fishbone or fault tree rather than choosing one silently.
Blame checkDoes any row name an individual or read "person did not do X"? If yes, the chain is incomplete: add the two missing whys (why was it not done, and was it possible to do) before proceeding.
Effectiveness reviewDate the action will be verified as effective, the measure used, and who confirms it. An unverified corrective action is an intention, not a control.

Two notes on using it. Keep the whole thing on one page; a 5 Whys that needs a second page has become something else and should be moved to a proper investigation format. And attach it to the work order rather than filing it separately, because an RCA that lives outside the maintenance system of record will not be found the next time the same asset fails, which is the one moment it has any value. The corrective and preventive action loop that closes around this output is covered in CAPA: corrective and preventive action explained.

9. Failure modes of the method itself

The 5 Whys fails in a small number of predictable ways. This table is the diagnostic I would hand to anyone reviewing other people's completed forms.

What goes wrongHow to spot itWhat to do
Vague problem statementNo asset tag, no time, no consequence. Reads like a category ("pump failure").Rewrite the statement before touching the whys. Object, observed condition, when, how long, what it cost, any pattern.
Parallel list, not a chainThe chain does not read upwards with "therefore". Each row answers the original problem.Re-ask each why against the previous answer only. Discard rows that do not follow.
Silent branchingOne answer plainly has more than one cause, but only one appears. Different investigators produce different chains from the same facts.List the branches explicitly, then escalate to a fishbone or fault tree. Do not choose by convenience.
Terminates on a personFinal row names an individual or says "did not follow the procedure". Action is "retrain" or "toolbox talk".Add the two missing whys: why was it not followed, and was it followable. Re-derive the action from those.
Unevidenced stepsNo evidence column, or every row says "team discussion".Add the column and grade each row verified / inferred / assumed. Treat assumed rows as open investigation tasks.
Padded to fiveLevels four and five are abstract ("poor culture", "lack of resources") after a concrete level three.Delete the padding. Close the chain at the last controllable, evidenced step and say so.
Stopped too earlyFinal row is still a symptom nobody can own an action against.Keep asking. Six or seven whys is legitimate.
Ran too farFinal row is outside the organisation's control: supplier quality, market conditions, original capital budget.Step back to the last row you can act on. Keep the outer cause as context, not as the finding.
Hindsight and single-investigator biasWritten by one person, after the fact, and every step looks obvious. No alternative explanation was ever considered.Have a second person re-derive the chain independently from the same facts. Divergence is the signal to escalate.
No effectiveness checkAction closed the day it was assigned. No measure, no verification date.Define what "effective" looks like and a date to confirm it. Recurrence after a closed RCA is itself a finding.
Used on the wrong class of problemApplied to a multi-discipline, high-consequence or already-recurring failure.Use a method built for it. The 5 Whys is a triage and simple-failure tool, not an incident investigation method.

10. Where it lives in a working system

Two practical points about fitting the method into day-to-day operations, kept deliberately generic.

First, decide the trigger. A 5 Whys on every work order is theatre and will be filled in without thought within a month. Trigger it on something meaningful: a repeat failure on the same asset within a defined window, a breakdown on an asset above a criticality threshold, any failure that caused a service interruption, or any corrective work order above a cost or downtime limit. A small number of properly run analyses beats universal compliance with an empty form every time.

Second, on tooling: any maintenance system can hold this. A structured long-description field on the work order with the template pasted in is enough to start, and is better than a separate spreadsheet nobody can find. What matters is only that the analysis is attached to the asset history, is searchable when the asset fails again, and carries the corrective action as a real, owned, dated task rather than a sentence. If you are earlier than that in your maintenance systems journey, the ground-level view is in what is a CMMS, and for the quantitative end of failure analysis where the 5 Whys hands over to modelling, see from FMEA to RUL.

The idea to walk away with

The 5 Whys earns its place because it is the cheapest useful RCA method in existence and the only one a team will genuinely run within an hour of a failure, at the asset, without a facilitator. That accessibility is its entire value, and it is worth protecting.

Protecting it means three things. Interrogate the previous answer, not the original problem, so that what you produce is a chain and not a list. Demand evidence for every link, so that what you produce is an analysis and not a story. And recognise the branch point, so that you escalate to a fishbone or a fault tree instead of quietly discarding two thirds of the causes to keep the format tidy. A chain that ends on "the technician did not follow the procedure" has failed all three tests at once, which is why that ending is the fastest way to audit an organisation's RCA maturity.

Final thoughts

There is a version of this guide that would tell you the 5 Whys is simplistic and you should use a proper method. I do not think that is right, and I do not think it is useful. Most failures in most facilities are not complex, and for those a well-run 5 Whys with an evidence column will find a fixable cause in twenty minutes. The problem is not that the method is too weak. The problem is that it is usually run without a specific problem statement, without evidence, without a branch check, and with a form that quietly rewards blaming whoever was holding the spanner.

Fix those four things and you will get more out of the 5 Whys than most organisations get out of far heavier methodologies, at a fraction of the cost. Then, for the minority of failures that genuinely branch, be honest early and reach for a tool built for trees. The skill is not the asking of the whys. It is recognising, somewhere around the second one, which of the two situations you are actually in.

Disclosure

Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.

Want your RCA process to produce findings people can act on?

Independent advisory on root cause analysis practice, failure coding, corrective action workflow and the maintenance data that makes any of it verifiable. 22+ years across utilities, oil and gas, manufacturing, government and facility operations.

Book a conversation

Related reading: Root cause analysis methods: a step-by-step guide, Fishbone diagram for root cause analysis, Root cause analysis tools: 8 methods explained, Fault tree analysis (FTA), RCA examples: equipment failures, CAPA explained, Incident investigation, Failure codes: problem, cause, action. External reference: IEC (IEC 62740:2015 Root cause analysis).

Muhammad Abbas

CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.

Work with me
MAbbaz.com
© MAbbaz.com