Walk into almost any post-incident review in a plant, a hospital estate or a utility control room, and within twenty minutes somebody will draw a horizontal arrow with a box on the right and start hanging branches off it. That is a fishbone diagram, and in the hands of a disciplined facilitator it is one of the best thinking tools available to a maintenance team. In the hands of an undisciplined one it produces a photographed whiteboard, a warm feeling of thoroughness, and no change whatsoever to the failure rate. The difference between those two outcomes is not the diagram. It is what happens in the hour after the diagram is finished.
The message up front: a fishbone diagram generates hypotheses, not conclusions. Its entire value is breadth, forcing a team to consider contributing causes it would otherwise skip past. It carries no mechanism at all for deciding which of those causes is real. If you stop at the finished diagram you have produced a structured wall of plausible guesses. The work that makes it root cause analysis is testing the candidate branches against evidence and converging on the ones the evidence supports.
1. What a fishbone diagram is, and the three names it travels under
A fishbone diagram is a structured way of laying out the possible causes of a single, defined problem so that they can be seen together rather than argued one at a time. You write the problem in a box on the right, draw a horizontal spine running into it, and then draw diagonal bones off that spine. Each bone is a category of cause. Off each bone you hang the specific causes that fall into that category, and off those you can hang sub-causes. The finished drawing resembles a fish skeleton, which is where the popular name comes from.
You will meet the same tool under three names, and they are all the same thing:
- Fishbone diagram: the descriptive name, taken from the shape. This is the term you will hear most in maintenance and facilities work.
- Ishikawa diagram: the attribution name, after Kaoru Ishikawa, the Japanese quality engineer who popularised it. You will see this in quality management and academic literature.
- Cause and effect diagram: the functional name, describing what the tool does. This is the term that turns up most often in formal standards and process documentation.
There is no meaningful difference between them. If a consultant tells you an Ishikawa diagram is a more rigorous variant of a fishbone diagram, they are selling you a vocabulary lesson. Use whichever term your organisation already uses, because arguing about the name is the first of many ways this tool wastes time.
The origin matters slightly more than trivia usually does, because it explains the tool's shape. Ishikawa developed and popularised the method within Japanese post-war quality practice, in a context that assumed shop-floor operators, not specialists, would be doing the analysis. It was designed to be drawable on paper by a group of people with no statistical training, as one of a small set of accessible quality tools. That design intent is why it is so easy to use and why it is so easy to misuse: it deliberately has no gate that stops you writing down something untrue.
2. Where the fishbone sits in the formal landscape
A point worth getting right, because it is commonly stated backwards. No standard prescribes how you draw a fishbone diagram, how many bones it should have, or what categories to use. There is no certifiable fishbone method and no committee that will tell you your diagram is non-compliant.
But it does not follow that the technique is informal folklore outside the standards world. IEC publishes IEC 62740:2015 "Root cause analysis (RCA)", which sets out principles and process steps for after-the-event analysis and describes a set of named recognised techniques. The Fishbone/Ishikawa diagram is one of those described techniques, as is the "Why" method that most people know as 5 Whys, alongside causes-tree, events-and-causal-factors charting and fault tree analysis. So the accurate statement is: no standard prescribes how to run a fishbone, but IEC 62740:2015 describes it as one of several recognised RCA techniques. That is a materially different claim from "the fishbone has no standard", and the difference matters if you are writing a procedure that an auditor will read.
If you are documenting a wider technique-selection rationale, IEC 31010:2019 "Risk management - Risk assessment techniques" is the broader catalogue of assessment methods and is worth naming. Note the designation carefully: it is IEC 31010, not ISO 31010 and not ISO/IEC 31010:2019. Both documents are paywalled, so buy the current edition if you need the normative detail rather than working from a summary, including this one.
Why this framing helps you
Knowing that an international standard recognises the technique lets you defend using it in a formal investigation procedure. Knowing that the standard does not prescribe how to draw it means you are free to design the session, the category set and the evidence step to suit your assets, and nobody can tell you your six bones should have been eight.
3. Why it beats a 5 Whys when the causes branch
This is the most practically useful thing to understand about the tool, and the thing that should decide which method you reach for.
A 5 Whys analysis is linear. You take the problem, ask why, take the answer, ask why again, and walk backwards down a single chain until you reach something you can act on. When the causal story genuinely is a chain, this is the better tool: it is faster, it is easier to explain, and its output is a short, defensible narrative. A bearing failed because it was starved of lubricant, because the grease point was missed, because the route sheet did not list it, because the asset was commissioned without being added to the lubrication schedule. One chain, clean, actionable.
The problem is that 5 Whys forces you to commit to one chain at the first step. At every "why" you pick an answer and abandon the alternatives. That is an enormous amount of unexamined judgement in the hands of whoever answers first, which in most rooms is the most senior or the most confident person present. If the true cause structure is several contributing factors converging, a linear method will find one of them, document it convincingly, and miss the rest entirely. You then fix a real cause, the failure recurs, and everybody loses confidence in the method.
The fishbone's single structural advantage is that it holds multiple contributing causes in view simultaneously. Nothing is discarded to make room for the next question. A lubrication gap, a spare of the wrong specification, an operator running the machine outside its duty envelope, and a measurement fault that masked the rising temperature can all sit on the diagram at the same time, on different bones, without the team having to decide yet which one matters. For failures with several plausible contributors, and most real equipment failures have several, that is exactly the property you need.
The practical rule I would give a maintenance team: if the room can agree on the causal chain within a few minutes, use 5 Whys and get on with it. If the room is producing competing explanations, or if the failure has recurred after a previous single-cause fix, draw a fishbone first to get the full candidate set on the wall, then drill the chosen branch with 5 Whys. The two tools are complements, not rivals, and the broader selection question across the full method set is covered in the root cause analysis methods pillar.
4. The category sets, and which to use
The bones of the diagram are category headings, and several standard sets circulate. The classic set for manufacturing and equipment work is the 6Ms, which maps well onto how physical assets actually fail:
- Machine: the equipment itself. Design, condition, wear, installed alignment, protective devices, spares fitted, modifications.
- Method: the procedure. The written work instruction, the PM task content, the isolation procedure, the commissioning routine, the handover.
- Material: what goes in or on. Spares, consumables, lubricants, filters, seals, process media, fuel, water quality.
- Manpower (People): competence, staffing level, training currency, shift handover, supervision, fatigue. Many organisations now write this bone as "People", which is both less awkward and less gendered.
- Measurement: how you know what you know. Instrument calibration, sensor placement, alarm setpoints, data capture in the maintenance system, gauge accuracy. This bone is underused and it is often where the real story is.
- Mother Nature (Environment): ambient conditions and external factors. Temperature, humidity, dust, salt-laden air, vibration from adjacent plant, flooding, power supply quality.
For service and administrative problems, where there may be no machine at all, other sets fit better. A helpdesk failure, a permit process breakdown or a billing error does not decompose usefully into Machine and Material.
| Set | Categories | Best suited to |
|---|---|---|
| 6Ms | Machine, Method, Material, Manpower/People, Measurement, Mother Nature/Environment | Equipment failures, plant and manufacturing problems, utilities assets. The default for maintenance and reliability work. |
| 4Ms | Machine, Method, Material, Manpower | A trimmed 6Ms for a short session or a simple mechanical problem. Cheap to run but you lose the Measurement and Environment prompts, which is where subtle causes hide. |
| 4Ss | Surroundings, Suppliers, Systems, Skills | Service delivery problems: helpdesk performance, response times, contractor quality, soft services complaints. |
| 8Ps | Product, Price, Place, Promotion, People, Process, Physical evidence, Productivity/Quality | Service and commercial processes, customer-facing failures, administrative and back-office breakdowns. Overkill for a broken pump. |
| Custom set | Whatever your failure population actually clusters into | Mature teams with good failure history. Often the best option: derive the bones from your own recurring cause patterns rather than borrowing a generic set. |
Categories are scaffolding, not a taxonomy
The categories exist for one reason: to prompt a team to think in directions it would otherwise neglect. They are not a classification scheme to be defended. If the room spends fifteen minutes arguing whether "wrong grease specification issued by stores" belongs under Material or Method, the facilitator has lost control of the session. Write it on either bone, or on both, and move on. Nothing downstream in the analysis depends on which bone a cause was filed under, and a team that treats the category set as an ontology will produce a tidier diagram and a worse investigation.
5. How to actually run the session
Most of the value, and most of the failure, is in facilitation rather than drawing. The sequence I would recommend:
- Write a precise problem statement in the head first, and do not proceed until the room accepts it. "Pump problems" is not a problem statement. "Chilled water pump CHWP-03 tripped on motor overload four times between 12 and 26 of the month, each trip within 40 minutes of start-up" is. A vague head produces a diagram where every bone is plausible and nothing is testable, because there is no specific event for evidence to be relevant to. Time, place, asset, symptom, frequency. If you cannot write that sentence, stop and gather facts before booking the workshop.
- Generate causes in silence before anybody speaks. This single change improves fishbone sessions more than anything else. Give everyone five minutes and a pad, or sticky notes, and have them write candidate causes alone. Then collect. If you open with an open discussion instead, the first two contributions anchor the whole session, the team converges on the loudest or most senior voice's theory, and the technician who actually saw the failure says nothing. Silent generation first is not a team-building gimmick, it is a defence against anchoring.
- Then group onto the bones. Take the collected causes and place them under categories. Duplicates are useful information: a cause that three people wrote independently deserves attention. Do not let the grouping become a debate about definitions.
- Ask "why" once or twice within each branch. Push each cause one or two levels deeper to get sub-causes. "Bearing not lubricated" becomes "grease point not on route sheet" becomes "asset commissioned without lubrication schedule". This is where the fishbone and 5 Whys genuinely merge, and it is how you avoid a diagram of symptoms masquerading as causes.
- Then challenge every branch for evidence. Go round the diagram and ask one question of each candidate: what evidence would confirm or eliminate this, and do we have it? Mark each cause as supported, contradicted, or untested. This turns the diagram from a brainstorm output into a test plan, and it is the step that makes the rest worthwhile.
- Close with owners and dates on the tests, not on the fixes. The output of a fishbone session is a list of things to go and check, each with a name against it. Assigning corrective actions at this point, before the evidence step, is how organisations end up implementing four fixes for a single-cause failure.
On practicalities: keep the session under ninety minutes, cap the room at six to eight people, and make sure at least one person present physically attended the failure. A fishbone drawn entirely by office staff from a work order description produces office-staff causes.
6. The step everyone skips: a fishbone produces hypotheses
Say this plainly, because it is the whole argument of this article. A completed fishbone diagram is a list of hypotheses. It is not a finding, not a root cause, and not a conclusion.
The tool has no truth test built into it. Every cause on the diagram got there because somebody in the room thought it was plausible. Plausibility is a very low bar, which is precisely why the technique is good at breadth. A well-facilitated session on a single pump trip can easily produce thirty candidate causes across six bones, and by construction most of them are wrong. The diagram cannot tell you which ones. It was never designed to.
So the step that converts the exercise into root cause analysis is convergence: taking each candidate branch and testing it against evidence until you can support or eliminate it. The evidence is usually mundane and already in your possession.
- Maintenance history: was the PM actually completed and when, what did the last three work orders on this asset say, what parts were consumed. Your failure coding structure either makes this searchable or makes it useless, which is one more reason coding discipline matters.
- Trend and process data: the BMS, SCADA or historian trace either shows the temperature rising over three weeks or it does not. This settles a great many arguments in one screenshot.
- Physical inspection of the failed part: the wear pattern on a bearing, the witness marks on a shaft, the state of a filter. Physical evidence eliminates hypotheses faster than any meeting.
- Documents: the actual procedure as written, the actual spare part specification issued, the training record, the calibration certificate. Not what people believe those documents say.
- Interviews: what the operator and technician observed, asked as open questions about what happened rather than why it happened.
Some hypotheses cannot be tested with what you have. That is a legitimate finding: record them as untested, and if the failure recurs, they are your starting point. What is not legitimate is treating an untested hypothesis as a cause and raising corrective actions against it. That is how CAPA registers fill up with actions nobody can close and nobody believes in, which quietly destroys the credibility of the whole investigation process.
The test to apply before you close the analysis
For each cause you are about to act on, answer two questions. First: what specific evidence supports this, and where is it filed? Second: if this cause were eliminated, would the failure have been prevented or made materially less likely? A cause that survives both questions is worth a corrective action. A cause that survives neither is a guess with a box drawn round it.
7. A worked fishbone example
The following is a hypothetical illustration I have constructed for teaching purposes. The asset, the readings and the causes are invented; none of it is drawn from a client. I have rendered it as a structured table rather than a drawing, because in practice a table is easier to circulate, easier to keep in a document management system and easier to add an evidence column to, which is the column that matters.
Problem statement (the head of the fish): Hypothetical chilled water pump CHWP-03 tripped on motor overload on four occasions across a two-week period, each trip occurring within roughly 40 minutes of start-up. No trips recorded on the duty pair CHWP-01 or CHWP-02 in the same period.
| Bone | Candidate cause | Sub-cause (one why deeper) | Evidence to test it |
|---|---|---|---|
| Machine | Bearing degradation raising motor load | Bearing past service life, not replaced at overhaul | Vibration reading; strip and inspect bearing; overhaul record |
| Impeller partially obstructed | Strainer bypassed during previous repair | Open and inspect strainer and impeller; last repair work order | |
| Coupling misalignment after recent repair | No alignment check recorded on reinstatement | Laser alignment check; repair work order task list | |
| Method | Start-up sequence allows pump to run against closed valve | Sequence changed during controls upgrade, procedure not revised | BMS sequence-of-operation document; control logic as built |
| PM task list omits bearing lubrication | Asset added to register without inheriting the class PM template | PM template for pump class versus the task list on this asset | |
| Material | Incorrect grease specification applied | Substitute issued by stores when specified grade was out of stock | Stores issue record; grease sample; specification sheet |
| Replacement bearing not to original specification | Non-original part sourced on urgency | Part number on the fitted bearing versus the asset BOM | |
| People | Reinstatement carried out by technician unfamiliar with this pump class | Competence matrix not consulted when allocating the job | Work order assignment record; training and competence record |
| Operator manually selecting CHWP-03 as lead pump outside rotation | Rotation schedule not visible at the local panel | BMS operator action log; run-hour comparison across the three pumps | |
| Measurement | Overload relay set below correct trip value | Relay replaced and reset to a default rather than nameplate value | Relay setting versus motor nameplate full-load current |
| Rising current trend not visible to the team | Motor current not trended or alarmed in the BMS | Historian point list; whether the point exists and is logged | |
| Environment | Plant room ambient temperature elevated | Extract fan for the plant room failed and not reported | Plant room temperature log; fan status and work order history |
| Supply voltage imbalance across phases | Upstream distribution fault or loading change | Power quality log at the panel; thermographic survey of the board |
Thirteen candidate causes on six bones, from one pump tripping four times. Note two things about that table. First, almost all of it is plausible, and that is the point: had the team run a 5 Whys instead, it would have picked one of these, probably the bearing, and never written down the overload relay setting or the plant room extract fan. Second, every row has an evidence column, and most of those tests take an hour or less. The relay setting can be checked against the nameplate in ten minutes. The run-hour comparison is one query. That is the convergence step, and in this illustrative case it would very likely eliminate nine or ten of the thirteen rows in a single afternoon, leaving a small, defensible set of supported causes to act on.
If you want the same problem viewed through the lens of failure mode rather than cause category, the failure modes guide covers the complementary taxonomy, and it is a useful cross-check: a fishbone branch that does not map to any credible failure mode for that asset class is usually one to eliminate early.
8. How it combines with other tools
The fishbone is at its best as the opening move in a sequence rather than a standalone exercise. The combinations that genuinely work:
- Fishbone then 5 Whys. Use the fishbone for breadth to get the candidate set, then take the branch the evidence supports and drill it with 5 Whys to reach a systemic cause you can actually fix. This is the highest-value pairing and the one I would default to. Breadth first, then depth on the branch that survives.
- Fishbone then Pareto. If you are analysing a recurring problem rather than a single event, use the fishbone to define the candidate cause categories, then count occurrences across your failure history and apply Pareto analysis to see which categories dominate. This converts subjective plausibility into frequency data, which is a genuine improvement. It requires that your history is coded well enough to count, which brings you back to failure coding discipline.
- Fishbone then fault tree. Where you need logical structure and probability rather than a list, fault tree analysis is the upgrade. A fault tree expresses how causes combine using AND and OR logic and can be quantified. Use the fishbone to populate the basic events, then build the tree. This is worth the effort on high-consequence safety or availability problems and is overkill for a routine equipment failure.
- Fishbone into CAPA. Whatever survives the evidence step feeds the corrective and preventive action process, and the discipline of separating a correction (fix this pump) from a preventive action (fix the PM template for the whole pump class) is where the durable value sits. The CAPA guide covers that split.
For the wider set of techniques and when each earns its place, the RCA tools comparison sets the fishbone against the alternatives rather than treating it as the default.
9. How the method fails
The fishbone fails in a small number of very predictable ways. Every one of them is a facilitation or follow-through failure rather than a flaw in the tool.
| Failure mode | What it looks like | The fix |
|---|---|---|
| Categories treated as the answer | The report concludes "the cause was Method and Manpower". A category is not a cause. Nobody can act on it. | Require every documented cause to be a specific, checkable statement about a specific thing. No corrective action may reference a bone name. |
| No evidence testing | The diagram is the deliverable. Every branch is treated as equally true, and actions are raised against all of them. | Make the evidence column mandatory. The session output is a test plan with owners, not an action list. |
| Photographed and never revisited | A phone picture of a whiteboard in a WhatsApp group. No owner, no transcription, no follow-up. | Transcribe into the document or maintenance system the same day, attached to the work order or incident record, with named owners and dates. |
| Diagram too large to act on | Eighty causes across nine bones. Technically thorough, operationally paralysing, so nothing happens. | Cap the session. After grouping, have the team vote or rank to a shortlist of the most credible candidates and test those first. |
| Anchoring on the loudest voice | The diagram documents the supervisor's opening theory in six categories. Dissenting views were never written down. | Silent individual generation before any discussion. Always. |
| Vague problem statement | The head says "high downtime" or "poor reliability". Every conceivable cause fits and none can be tested. | Refuse to start without a specific event or defined, measured problem. Send the room away to gather facts if needed. |
| Blame instead of cause | The People bone fills with named individuals, everybody becomes defensive, and the useful information stops flowing. | Write system causes, not names: absent procedure, missing competence check, unclear allocation. Note that IEC 62740:2015 explicitly excludes assigning blame from the scope of root cause analysis. |
| Wrong tool for the problem | A fishbone drawn for a problem the team already understands, purely because the procedure says to do one. | If the chain is obvious, use 5 Whys. If you need probability and logic, use a fault tree. Method selection is a real decision. |
Where I would not use a fishbone at all
Three cases. When the causal chain is already clear to everyone in the room, because the diagram adds ceremony and no insight. When you need to quantify how causes combine, because the fishbone has no logic operators and no probabilities, and a fault tree does. And when the problem is a pattern across hundreds of events rather than a single failure, because there the right first move is to count and rank the history, not to brainstorm. The fishbone is a tool for one defined problem with several plausible contributors and a team that disagrees about which. Outside that, something else is usually better.
10. Making it stick in the maintenance system
A fishbone that lives only on a whiteboard has a half-life of about two days, so the last practical question is where the output goes. Some advice that costs nothing:
- Attach it to the record, not to a folder. The analysis belongs against the work order or the incident record for the failure, wherever your team already looks. Filed in a shared drive by date, it will never be found again when the same pump trips next year.
- Feed the supported causes back into failure coding. If the analysis found a cause your code list cannot express, that is a gap in the code list. Over a few years this is how a genuinely useful cause taxonomy gets built, and it is what makes future Pareto analysis possible at all.
- Convert systemic causes into register changes, not just repairs. If the cause was a PM template missing a task, the action is to fix the template for the asset class. One analysis correctly followed through can prevent a category of failures rather than one instance.
- Watch what the backlog tells you. Recurring failures on the same asset, visible in backlog and downtime tracking, are the signal that a previous single-cause analysis was incomplete and a fishbone is warranted.
- Keep the tooling boring. A table in a document, or a structured note attached to the work order, is genuinely sufficient. Most maintenance systems will store the analysis as an attachment or a long text field, and that is fine; a CMMS is where the record should live so that the history is searchable later. Dedicated RCA software is not the bottleneck. The bottleneck is whether anyone runs the evidence tests.
The idea to walk away with
The fishbone diagram does one thing that no linear method can do: it holds several plausible contributing causes in view at the same time, so a team stops collapsing onto the first explanation offered by the most confident person in the room. That is a real and valuable property, and it is why the technique has survived seventy years and earned a place among the recognised RCA techniques described in IEC 62740:2015.
But breadth is all it gives you. It has no mechanism for deciding which of those causes is true, and it never claimed to. The finished diagram is the halfway point, not the deliverable. The analysis happens afterwards, in the unglamorous work of going and checking the relay setting, pulling the stores issue record, comparing the run hours and stripping the bearing, until the plausible set has been reduced to the supported set. Do that and the fishbone is one of the best tools you have. Skip it and you have drawn a fish.
Final thoughts
If I had to leave a maintenance team with one habit from this article, it would be the silent generation step at the start and the evidence column at the end. Those two changes cost nothing, require no software, need no training course, and between them fix most of what goes wrong with the method. Everything else is detail: which category set you use, how many bones you draw, whether you call it a fishbone or an Ishikawa diagram. None of that determines whether the exercise was worth the room.
And be honest about method selection. A procedure that mandates a fishbone for every incident will produce a lot of compliant diagrams and not much learning. Draw one when the causes genuinely branch and the room genuinely disagrees. When the chain is obvious, say so, fix it, and give the hour back to the team.
Disclosure
Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.
Building a root cause analysis process that holds up?
Independent advisory on failure investigation practice, cause coding structures, CAPA follow-through and getting the analysis into the maintenance system where it can be found again. 22+ years across utilities, oil and gas, manufacturing, government and facility operations.
Book a conversationRelated reading: Root cause analysis methods: a step-by-step guide, The 5 Whys method and examples, RCA tools: 8 methods explained, Pareto analysis, Fault tree analysis, Failure modes and how to analyse them, CAPA explained, Failure codes: Problem, Cause, Action.
Muhammad Abbas
CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.
Work with me