Most reliability methods start at the bottom. You take a component, list the ways it can fail, and reason forward to what each failure would do. Fault tree analysis does the opposite. It starts with one precisely defined undesired event, the thing that must not happen, and works backwards asking a single repeated question: what combination of conditions would have to be true for this to occur? That change of direction is not a stylistic preference. It is what makes FTA capable of finding failure paths that depend on several ordinary conditions coinciding, which is exactly the class of failure that surprises well run organisations.
The message up front: the value of a fault tree is not the drawing. It is the list of minimal cut sets you read off it, the smallest combinations of basic events that are sufficient to cause the top event. Any cut set containing exactly one element is a single point of failure, stated in writing, with the logic that proves it. That finding changes designs, specifications and protection schemes. Everything else in the method exists to produce it.
1. What fault tree analysis actually is
Fault tree analysis is a top-down, deductive logic method. You define one undesired outcome, call it the top event, and then decompose it downwards through logic gates into the events and conditions that could bring it about. Each layer answers "what has to happen for the layer above it to happen", and you keep going until you reach events you are willing to treat as indivisible.
The word deductive is doing real work there. Deductive reasoning goes from a general statement to the specific conditions that would satisfy it, so FTA reasons from a defined system state backwards to its possible causes. Inductive reasoning goes the other way, which is what failure mode and effects analysis does when it starts at a component failure mode and asks what effect it has on the system. The two methods travel in opposite directions through the same territory, and that single fact explains almost every practical difference between them.
The method originated in aerospace and nuclear safety work and is formalised internationally as IEC 61025:2006 (Edition 2) , "Fault tree analysis (FTA)". That is a voluntary, paywalled international standard, and it is still the current edition, which means the formal description of the method is now twenty years old. I do not read that as neglect. FTA is a mature technique whose logic has not needed revision. But it does mean the standard predates most of the modelling software and all of the data availability arguments that now dominate practical FTA work, so treat it as the reference for the logic rather than for tooling.
2. Defining the top event, the step that decides everything
If you get one thing right in a fault tree, make it the top event. It determines the boundary of the analysis, the level of detail that is appropriate, and whether the finished tree tells you anything you can act on. A loose top event produces a tree that sprawls without conclusion, and teams usually discover this the hard way, three days into drawing, when somebody asks what the tree is for. A usable top event is specific about the failure state, the system boundary, the operating condition and the consequence level. Compare two attempts on the same system:
- Too loose: "Cooling system failure." Which cooling system? Failure of what function? Under what load? Partial or total? This will generate a tree with a hundred branches and no conclusion, because every conceivable cooling defect belongs somewhere under it.
- Usable: "Total loss of chilled water supply to Data Hall B for more than five minutes while the hall is at design IT load." Now the boundary is fixed, the function is named, the threshold is measurable, and the operating condition is stated. Anything that does not contribute to that specific state is out of scope, which is what makes the tree finite.
Two habits help. Write the top event as a sentence and get it agreed by operations before drawing anything, because they will tell you whether five minutes is the right threshold. And derive it from an existing criticality assessment rather than inventing it, so the analysis points at something the organisation has already agreed matters. The ranking work in the equipment criticality analysis guide is the natural feeder.
The test for a good top event
Read the top event aloud and ask whether two competent engineers would draw the same first layer of the tree underneath it. If they would not, the top event is still ambiguous. Tighten it before you go further, because every hour spent below an ambiguous top event is wasted.
3. The logic gates and what they mean physically
Gates are the grammar of a fault tree. They state how the events below a gate combine to produce the event above it. In practice the overwhelming majority of real trees are built from two gates, AND and OR, and understanding what those two mean physically is most of the value of the method.
An OR gate says the event above occurs if any one of the events below occurs. An AND gate says the event above occurs only if all of the events below occur together. Now translate that into engineering:
- An OR gate is where single points of failure live. If a branch sits under an OR gate on its own, that one event is sufficient to drive the outcome above it. Trace a path of OR gates from a basic event all the way to the top event and you have found a single component whose failure takes down the whole function.
- An AND gate is where redundancy lives. An AND gate is the logical expression of "both have to fail". Every genuine redundancy in the design should show up as an AND gate somewhere in the tree. If your two-pump arrangement appears under an OR gate, either you drew it wrong or the redundancy is not real.
That observation alone justifies the exercise. Walk a finished tree with the operations team and point at the gates: here you are protected, and here one item stands between you and the top event. That conversation can change a design where a formal report had failed to.
| Gate | Logic | What it means physically |
|---|---|---|
| OR | Output occurs if any input occurs | No protection at this point. Each input is sufficient on its own. A chain of OR gates up to the top event marks a single point of failure. |
| AND | Output occurs only if all inputs occur | Redundancy, diversity or a protective layer. All the inputs must be defeated together, so the combination is what matters, not any single item. |
| Voting (k out of n) | Output occurs if at least k of n inputs occur | Partial redundancy. Three chillers where any two carry the load, or a two out of three instrument trip arrangement. Sits between OR and AND. |
| Inhibit / conditional | Output occurs if the input occurs while an enabling condition holds | Failures that only matter in a particular state: during startup, at design load, when the changeover is in manual. Useful, and frequently abused as a way of hiding assumptions. |
| Priority and exclusive variants | Order or mutual exclusivity matters | Sequence dependent failures, where B only causes the outcome if A has already happened. Other gate types exist beyond these; if you find yourself needing them often, the tree is probably modelling a dynamic behaviour better handled elsewhere. |
For maintenance and facilities work, stay with AND, OR and voting gates unless there is a compelling reason not to. The exotic gate types are legitimate, but a tree that uses them liberally becomes unreadable to the people who have to act on it.
4. Basic, intermediate and undeveloped events, and where to stop
Fault trees distinguish a handful of event types, and the distinctions are practical rather than academic.
- The top event is the single undesired outcome at the head of the tree.
- Intermediate events are the failure states in the middle layers. They are the output of a gate and the input to the gate above. They exist to make the logic readable and to give the tree its structure.
- Basic events are the leaves. A basic event is a failure you have chosen not to decompose further, either because it is genuinely elementary at your level of analysis or because decomposing it would add no useful information. "Pump P-01 fails to start on demand" is a perfectly good basic event for a facilities tree even though a reliability engineer could break it down into bearing, winding and contactor failures.
- Undeveloped events are branches you deliberately stop developing and say so. Usually because the detail sits outside your boundary, the information is not available, or the branch is known to be insignificant relative to the others. The honest practice is to mark them as undeveloped rather than quietly leave them as basic events, so a reader can see where the analysis chose to stop.
Knowing where to stop separates a useful tree from an exercise. The rule I would give: stop decomposing when further detail would not change a decision. If breaking a pump into components does not change what you would do about the finding, do not break it down. If the tree informs a spares or task decision, component level may genuinely matter, and that is when you go deeper. For the vocabulary of component level failure behaviour underneath a basic event, the failure modes guide covers the ground.
5. Building a tree: a worked hypothetical
Here is a small tree on an invented system so the logic is visible. It is entirely hypothetical: a chilled water plant serving one critical hall, with two chillers, two pumps on a common header, and a single isolation arrangement. No real site is described and no probability or failure rate values appear in it. Top event: total loss of chilled water supply to Hall B for more than five minutes at design load. Working down one layer at a time, the question at each gate is what would have to be true for the event above to occur.
| Tree structure (indentation shows depth) | Event type | Gate below |
|---|---|---|
| T · Total loss of chilled water to Hall B > 5 min at design load | Top event | OR |
| I1 · No chilled water generated | Intermediate | OR |
| I1.1 · Both chillers unavailable | Intermediate | AND |
| B1 · Chiller CH-01 fails or is out for maintenance | Basic | Leaf, no gate |
| B2 · Chiller CH-02 fails or is out for maintenance | Basic | Leaf, no gate |
| B3 · Loss of electrical supply to the whole plant room | Basic | Leaf, no gate |
| I2 · Chilled water generated but not delivered | Intermediate | OR |
| I2.1 · No flow from the pump set | Intermediate | AND |
| B4 · Pump P-01 fails to run on demand | Basic | Leaf, no gate |
| B5 · Pump P-02 fails to run on demand | Basic | Leaf, no gate |
| B6 · Common header isolation valve closed or seized shut | Basic | Leaf, no gate |
| B7 · Header rupture or major leak downstream of the pumps | Basic | Leaf, no gate |
| U1 · Building control system fails to command the plant | Undeveloped | Not developed: controls sit outside this boundary |
Three things are already visible from a tree this small, which is the point of building it. First, the redundancy on the chillers and on the pumps has correctly resolved into AND gates, so the design intent is confirmed. Second, three basic events sit under OR gates with nothing beside them: the plant room electrical supply, the header isolation valve, and the header itself. Third, one branch is honestly marked undeveloped rather than pretended away, so any reader knows the controls contribution has not been assessed.
6. Reading the tree: minimal cut sets are the real output
A cut set is any combination of basic events that, if they all occur, is sufficient to cause the top event. A minimal cut set is one you cannot remove anything from without breaking its sufficiency. Minimal cut sets are the actual product of a fault tree, and a tree that has never been reduced to its cut sets has not yet paid for itself.
Reading the hypothetical tree above, the minimal cut sets are:
- {B3} loss of electrical supply to the plant room, on its own.
- {B6} header isolation valve closed or seized shut, on its own.
- {B7} header rupture downstream of the pumps, on its own.
- {B1, B2} both chillers unavailable together.
- {B4, B5} both pumps failing to run together.
Three of the five cut sets contain a single element, and each is a single point of failure the tree has not merely asserted but demonstrated. The two chillers and two pumps that the design brief was proud of contribute two cut sets of size two, while a single valve and a single length of header contribute cut sets of size one. That is the finding that changes a design, and you cannot get it from a bottom-up method without a great deal more work.
The reading rules are simple. Cut set size orders your concern before any probability enters the picture: size one is a single point of failure, size two means two things must coincide, larger sets are progressively less likely to align. A basic event appearing in several cut sets is a common cause candidate. And a cut set whose elements share a power source, a room, a maintenance crew or a control loop is not as independent as its size suggests, which is where common cause analysis sits alongside the tree.
The one output to insist on
If you commission a fault tree from anybody, internal or external, make the ranked minimal cut set list the deliverable, not the diagram. The diagram is working material. The cut set list is the engineering finding, and it is the thing that goes into a design change, a protection review or a spares decision.
7. Qualitative versus quantitative FTA
A qualitative fault tree stops at the structure and the cut sets. It tells you which combinations are sufficient, which are single points of failure, and where common causes cut across branches. A quantitative fault tree assigns a probability or failure rate to every basic event and propagates those values up through the gates to produce a probability for the top event.
Quantification is genuinely valuable where the input data exists. In nuclear, aerospace and parts of process industry it does, because those sectors spent decades building failure rate datasets under disciplined taxonomies. In most maintenance and facilities organisations it does not: failure history is patchy, coding is inconsistent, downtime is not reliably captured, and the population of identical items is too small to estimate a rate from. Multiplying numbers you do not have produces a top event figure with the appearance of precision and none of the substance.
So my position is that qualitative FTA is where most organisations should stop, and stopping there is not a compromise. The cut set list does not need probabilities to be actionable. A single point of failure on a critical supply is a finding regardless of its rate.
If you do quantify, label every input value as what it is: measured from your own history, taken from a named published dataset, or engineering judgement. Carry that labelling through to the result so the reader knows how much of the top event figure was invented. An illustration: assign clearly invented values of 0.02 and 0.05 to two independent basic events under an AND gate purely to demonstrate the arithmetic, and the combination is 0.001. The only honest caption for that figure is that both inputs were made up to show how an AND gate combines, which is exactly the standard I would hold any real analysis to.
Improving the underlying data is the prerequisite, not an optional extra. The problem, cause and action coding in the failure codes pillar is what eventually makes quantification defensible, and the modelling relationship between life estimates and failure mode work is covered in from RUL to FMEA.
8. FTA versus FMEA, since readers conflate them
This is the comparison that causes the most confusion, and the confusion is understandable because both methods analyse failure and both produce a structured document. The difference is direction, and everything else follows from it.
| Dimension | Fault tree analysis (FTA) | Failure mode and effects analysis (FMEA) |
|---|---|---|
| Direction | Top down, deductive | Bottom up, inductive |
| Starting point | One defined undesired top event | Every component or function, and its failure modes |
| Core question | What combinations could cause this outcome? | If this fails in this way, what is the effect? |
| Handles combinations | Yes. This is its main strength: AND gates and multi element cut sets are native to the method | Poorly. It is fundamentally a single failure at a time analysis |
| Coverage | Deep but narrow. Only what can cause the chosen top event | Broad but shallow per item. Every item gets considered |
| Primary output | Minimal cut sets and identified single points of failure | A ranked register of failure modes with effects, controls and actions |
| Best at | Testing redundancy claims, protection and interlock logic, investigating an event where conditions coincided | Task selection, systematic design review, driving a maintenance programme from failure modes |
| Typical effort profile | Concentrated: one top event, a small team, days rather than months | Extensive: item by item across a whole system, sustained over a long period |
| International reference | IEC 61025:2006 (Edition 2), "Fault tree analysis (FTA)" | IEC 60812:2018 (Edition 3), "Failure modes and effects analysis (FMEA and FMECA)". The title changed at Edition 3, so older citations of it look different |
The methods are complementary, not competing, and the sequence that works in practice is to use FMEA to understand how things fail and FTA to understand how the system as a whole can be defeated. FMEA gives you a vocabulary of failure modes at component level; those become well informed basic events in a fault tree. FTA then tells you which of those modes matter in combination, which FMEA structurally cannot. Neither replaces the other, and an organisation with only one of them has a blind spot in a predictable place. The full method treatment sits in the FMEA guide, and the decision logic that consumes both sits in reliability centred maintenance.
9. Where FTA genuinely earns its cost
Fault tree analysis is not cheap. It needs a competent facilitator, engineers who know the system, and enough time to argue about gates. That cost is justified in a fairly narrow set of situations, and being honest about which ones saves a lot of wasted effort.
- Safety critical systems where the top event is severe. When the undesired outcome involves injury, environmental release or major asset loss, explicit logic is proportionate to the stake. In the United States, and only there, fault tree is one of the methodologies named as acceptable for a process hazard analysis under the process safety management rule at 29 CFR 1910.119(e) . That is a federal US requirement with no force elsewhere, and note that a process hazard analysis works at process and scenario level. It is not a task level analysis and should not be confused with one.
- Protection, interlock and trip logic. Protection schemes are pure logic, which makes them the most natural possible subject for a logic method. A fault tree on a trip arrangement will find the case where the protection itself is the single point of failure, which reading the cause and effect matrix rarely does.
- Testing a redundancy claim. This is where the fastest return tends to appear in facilities and utilities work. Somebody has specified N plus one and believes the function is protected. A fault tree either confirms that as an AND gate or exposes a shared valve, shared board, shared control loop or shared room that collapses the redundancy into an OR. It is a short piece of work with a disproportionate finding.
- Investigating a serious event where several conditions had to coincide. Linear cause techniques struggle here by design. When an incident required a guard to be defeated and an alarm to be in a suppressed state and a relief path to be isolated, you need a method that represents combination explicitly, and that is a fault tree. For the wider selection of investigation methods and where each fits, see the root cause analysis methods pillar and the practical survey in eight RCA tools explained.
Conversely, FTA is the wrong tool for routine single cause failures. If a belt snapped and the reason is that nobody inspected it, run five whys and move on. Reaching for a fault tree on a simple failure signals that the method has become an end in itself.
10. Where FTA sits in the standards landscape
A short orientation, because readers often ask what document governs this and get a confused answer.
- IEC 61025:2006 (Edition 2), "Fault tree analysis (FTA)" is the dedicated international standard for the method itself. Voluntary and paywalled, and still the current edition after twenty years.
- IEC 31010:2019, "Risk management - Risk assessment techniques" catalogues risk assessment techniques including fault tree analysis, and is the useful reference when you are choosing between techniques rather than executing one. Note the designation carefully: it is IEC 31010, not ISO 31010, and not a dual prefixed designation. Only the withdrawn 2009 edition carried a joint prefix, and the error is repeated constantly.
- IEC 62740:2015, "Root cause analysis (RCA)" is a real international standard for root cause analysis, and it describes named RCA techniques including fault tree approaches alongside others. Anyone who tells you root cause analysis has no standard is mistaken. IEC 62740 covers after the event analysis and explicitly stays away from assigning blame.
- IEC 60812:2018 (Edition 3), "Failure modes and effects analysis (FMEA and FMECA)" is the companion document for the bottom up method in the comparison above. Its title changed at Edition 3, which is why older references to it read differently.
- Bowtie comes up in the same conversations and is worth placing correctly. It is a 2018 concept book from CCPS and the Energy Institute, not a standard. It appears as a listed technique in IEC 31010:2019, but there is no bowtie standard, and calling it one misleads readers about its status.
None of these documents is law anywhere by itself. They bind through contracts, client specifications and regulator expectations. I am describing what they cover rather than quoting them, because they are paywalled. If a specification requires you to work to one of them, buy the current edition and read it.
11. How fault tree analysis fails in practice
The method is sound. The ways it goes wrong are organisational and they are predictable.
- The tree is built too deep. A team enjoying the decomposition keeps going until every basic event is a fastener. The tree becomes unreadable, the cut set list becomes enormous, and the findings that mattered are buried. Depth is not rigour. Stop where further detail stops changing decisions.
- Gates are used loosely. The most damaging single error in FTA is an OR gate drawn where the logic is AND, or the reverse. An OR drawn as an AND hides a single point of failure, which is precisely the finding you built the tree to get. Every gate deserves the question "is that really all of them, or really any of them", asked out loud.
- Quantification with invented probabilities. The temptation to fill the tree with numbers is strong because a top event probability looks like a result. If the numbers came from nowhere, the probability came from nowhere, and a decision made on it is worse than a decision made on the qualitative cut sets alone. Say plainly when data is not available.
- Independence is assumed where it does not hold. An AND gate over two items on the same board, maintained by the same crew on the same day, is not the protection the gate implies. Examine common cause explicitly or the tree flatters the design.
- Nobody maintains the tree after a design change. This is the most common quiet failure. A tree is built during a project, the findings are actioned, and then two years of modifications go by. The tree still hangs in the plant room describing a system that no longer exists, and the next person to read it is misled by it. Either put the tree under change control with the drawings or take it down.
What FTA cannot do for you
A fault tree only knows what the team put into it. It will not surface a failure path nobody thought of, it does not represent timing and sequence well, it says nothing about degradation or how a failure develops, and it is a snapshot of one configuration at one moment. It is also silent on human and organisational factors unless they are modelled deliberately as events, and modelling those honestly is difficult. Treat a tree as a structured record of the team's understanding, not as an exhaustive statement of how the system can fail.
12. Running a fault tree that gets used
A practical sequence for a maintenance or engineering organisation attempting this for the first time:
- Pick a top event that already matters. Take it from the criticality assessment or from a recent serious event. Do not invent one to practise on, because the output will not be actioned and the exercise will not be repeated.
- Write the top event as a full sentence and get it agreed. Function, boundary, threshold, operating condition. Agreed by operations before drawing starts.
- Fix the boundary and write down the assumptions. What is in, what is out, what configuration the system is in, what support systems are assumed available. Undeveloped events go here.
- Build one layer at a time, deciding the gate before the events. Deciding the gate first forces the logic question and prevents the tree drifting into a cause list.
- Stop at the level where findings change decisions. Mark undeveloped branches honestly.
- Reduce to minimal cut sets and rank by size. Single element cut sets first. This is the deliverable.
- Test the cut sets for common cause. Shared power, shared space, shared controls, shared maintenance. A size two cut set with a shared dependency behaves like a size one.
- Convert findings into owned actions with dates. Design change, protection change, spares holding, procedure change, or an accepted and recorded risk. An unactioned cut set list is the same as no analysis.
- Put the tree under change control. Review it when the system changes, not on an annual calendar cycle that nobody honours.
On software: general risk and reliability tools will draw trees and compute cut sets, which is worth having for large trees because manual reduction becomes error prone quickly. For a tree of the size shown above, a whiteboard and a spreadsheet are adequate, and the whiteboard has the advantage that the arguments about gates happen in the room. Whatever you use, the cut set list and the assumptions belong somewhere durable and version controlled alongside the asset records.
The idea to walk away with
Fault tree analysis is a top down deductive method whose real product is a list of minimal cut sets. AND gates mark where you are protected, OR gates mark where you are not, and any cut set of size one is a single point of failure written down with the logic that proves it. That is the finding that changes designs, and you can get to it without a single probability figure.
It complements FMEA rather than replacing it. FMEA tells you how components fail; FTA tells you which of those failures matter in combination.
Final thoughts
The reason I keep recommending fault trees for redundancy claims specifically is that the exercise is short and the finding is frequently uncomfortable. It takes a few days to test whether an N plus one arrangement is genuinely N plus one, and often it is not, because of a valve, a board, a room or a control loop the block diagram never showed as shared. No amount of monitoring or scheduling compensates for that. It is a structural property of the design, and only logic finds it.
Keep trees shallow enough to read, be strict about the gates, be honest about the branches you did not develop, stay qualitative unless your data supports arithmetic, and put the finished tree under the same change control as the drawings. Do that and FTA is one of the highest value few days an engineering team can spend on a critical system.
Disclosure
Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.
Need a redundancy claim tested properly?
Independent facilitation of fault tree and failure mode analysis on critical systems, plus the failure coding and asset data work that makes the findings stick. 22+ years across utilities, oil and gas, manufacturing, government and facility operations.
Book a conversationRelated reading: Root cause analysis methods: a step by step guide, FMEA: failure mode and effects analysis, Root cause analysis tools: 8 methods explained, Failure modes: common types and how to analyse them, Equipment criticality analysis, Reliability centred maintenance introduction.
Muhammad Abbas
CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.
Work with me