Ask a maintenance team where their problems are and you will get opinions. Ask a Pareto chart and you get a ranked list in about twenty minutes. That is why Pareto analysis has survived a century of management fashion: it is arithmetic, it needs no licence, and it turns an argument into a picture. But the same tool has also sent teams chasing the wrong equipment for a year, because a chart built on failure counts had buried the one failure mode that mattered somewhere near the tail. The technique is sound. The way it is usually applied is not.
The message up front: a Pareto chart does not tell you what your biggest problem is. It tells you what your biggest problem is according to the measure you chose, as coded by whoever closed the work orders. Change the measure and the top category often changes. Change the coding discipline and the whole chart changes. Run it on at least two measures, always state which measure you used, and treat it as the tool that decides where to look, never the tool that decides what to do.
1. What Pareto analysis actually is
Pareto analysis is a prioritisation technique built on a single observation: when you break an outcome down by its contributing causes, the contributions are usually uneven. A handful of categories account for a large share of the total, and a long tail of categories accounts for very little. If that unevenness holds in your data, then work applied to the few large categories buys more improvement per hour than work spread evenly across all of them.
That is the whole idea. It is not a statistical test, it does not establish causation, and it produces no insight of its own. What it does is convert an unranked pile of records into a ranked list with a visible cut-off, so that a planning conversation can start from evidence instead of from whoever spoke loudest in the morning meeting. In maintenance that is a genuinely valuable thing to have, because maintenance backlogs are almost always longer than the resource available to clear them, so the real question is never "what should we fix" but "what should we fix first".
The technique sits in the family of structured problem-solving tools, and it is usually the first one you reach for, before any causal analysis at all. It belongs upstream of the methods described in the root cause analysis pillar: Pareto picks the target, the causal tools then work on it.
2. The 80/20 rule, treated honestly
Almost every introduction to this topic opens with the 80/20 rule: eighty percent of the effects come from twenty percent of the causes. The phrasing traces back to an early twentieth century observation about the distribution of wealth, later popularised as a general management heuristic and attached to the Pareto name. It is a useful mental shorthand. It is not a law of nature, it is not a property of maintenance data, and it is not something you should ever assert about a specific plant before you have looked.
Here is the honest position. The underlying observation is robust: real-world contributions are very often unevenly distributed rather than flat. The specific 80/20 split is a rule of thumb that happens to be memorable. When you actually compute it on your own work order history, the split might be closer to 90/10, or 70/30, or, on a well-managed estate with a homogeneous asset base, something as flat as 60/40. All of those are legitimate findings. None of them is a failure of the technique.
So treat anyone who quotes 80/20 as a fixed property of your data with suspicion, including a vendor whose dashboard has an "80% line" drawn on it by default. The correct use is to compute the actual cumulative curve from your own records and read the cut-off off the curve, wherever it happens to fall. If the curve turns out to be almost a straight diagonal, that is important information too: it tells you there is no small set of dominant causes, and Pareto is the wrong tool for this dataset.
The test I would apply
Do not report "80/20 confirmed". Report the number you actually found, in the form: "the top three of our nineteen failure categories account for X percent of total downtime hours in the last twelve months". That sentence is defensible, specific to your estate, and states its measure. "The 80/20 rule applies to our failures" is none of those things.
3. How to build a Pareto chart, step by step
The method is short enough to run by hand, and I would recommend running it by hand once before you let a reporting tool do it for you, because the manual version forces you to make the choices explicit.
- Step 1: define the effect you are trying to reduce. Unplanned downtime hours. Emergency work orders. Maintenance spend. Repeat callouts. Write it down as a single sentence, because a chart built on a vague effect cannot be interpreted later.
- Step 2: choose the categorisation. This is the dimension you will rank by: failure cause code, asset class, location, equipment tag, trade, shift. One dimension per chart. Mixing dimensions is the commonest construction error and it produces double counting.
- Step 3: choose the measure. Count of events, total downtime hours, total cost, or a weighted score. Section 4 is entirely about why this step deserves more thought than it usually gets.
- Step 4: set the window. A defined, stated period, long enough to include seasonal variation but recent enough to reflect the current asset base. Twelve rolling months is a reasonable default for most estates. State it on the chart.
- Step 5: aggregate. Count or sum your chosen measure per category. This is a group-by, nothing more.
- Step 6: sort descending. Largest contributor on the left. This step is what makes it a Pareto chart rather than a bar chart.
- Step 7: compute the cumulative percentage. Running total of the measure, divided by the grand total, expressed as a percentage, category by category from left to right. Plotted as a line across the bars, this is the part that tells you where the cut-off sits.
- Step 8: read the elbow, not the eighty percent mark. Look for the point where the cumulative line flattens. Categories to the left of the flattening are your candidates. Categories to the right are noise for now.
On presentation: bars descending on the left axis, cumulative line on a right axis scaled zero to one hundred percent, category labels legible, measure and period stated in the title. An "other" bar, if you must have one, always sits at the far right regardless of its size, because it is not a real category. If "other" is one of your biggest bars, stop and go and fix your coding before you interpret anything, for the reasons in section 5.
4. The choice of measure, which is where most maintenance Pareto charts go wrong
This is the section that matters. Everything else about Pareto analysis is arithmetic you cannot really get wrong. The choice of measure is a judgement, it is usually made by default rather than deliberately, and it determines the answer.
Ranking by failure count answers "what happens most often". Ranking by downtime hours answers "what costs us most availability". Ranking by cost answers "what consumes most budget". Ranking by a criticality-weighted impact answers "what hurts the business most". These are four different questions, and on a real estate they routinely produce four different top categories. A count-based chart is dominated by the frequent-but-trivial: door closers, lamp replacements, filter changes, nuisance alarms. The rare catastrophic event, the transformer fault or the chiller compressor failure that took a building offline for three days, appears once, and sits near the tail where nobody looks.
| Measure | What it surfaces | What it hides |
|---|---|---|
| Event count | High-frequency nuisance work, chronic repeat offenders, labour-consuming small jobs, callout volume drivers | Anything rare. A single catastrophic failure counts as one, identical in weight to one lamp change |
| Downtime hours | Availability killers, long-duration repairs, parts-lead-time problems, assets with no redundancy | Frequent short stoppages that never individually register, and failures on assets that are not in production service |
| Direct cost | Budget drivers: expensive spares, specialist contractors, expedited freight | In-house labour if it is not costed, consequential losses, and cheap failures with severe operational impact |
| Labour hours | Where your crew's capacity is genuinely going, which is often not where anyone assumed | Failures fixed quickly but at high material cost or high business impact |
| Criticality-weighted score | Business impact, safety and compliance exposure, single points of failure | Nothing structurally, but it inherits every argument in your weighting scheme, so it is the hardest to defend |
The practical guidance I would give any planner: run it on more than one measure and compare. It costs a second query. Where two measures agree on the top category, you have a strong candidate and an easy business case. Where they disagree, the disagreement is itself the finding, and it usually points at something real, such as a frequent failure mode that is cheap to fix and a rare one that is ruinous, both of which need attention but through completely different routes.
And whenever you present a Pareto chart, say which measure you used, in the title, out loud, in the email. "Top failure causes" is not a chart title. "Top failure causes by downtime hours, main plant, last twelve months" is. Many of the bad decisions that come off a Pareto chart come from someone reading a count-based chart as though it showed impact.
5. A worked example, and the same data re-ranked
The numbers below are my own hypothetical figures, invented purely to demonstrate the mechanics. They are illustrative only and should not be read as benchmarks, typical values, or anything drawn from a real estate. Assume a mid-sized facility, twelve rolling months, corrective work orders grouped by failure cause code.
First, ranked by event count, which is how the majority of out-of-the-box maintenance reports will show it:
| Rank | Failure category (illustrative) | Events | % of events | Cumulative % |
|---|---|---|---|---|
| 1 | Lighting and lamp failure | 210 | 35.0 | 35.0 |
| 2 | Door hardware and closers | 132 | 22.0 | 57.0 |
| 3 | AHU filter and belt issues | 96 | 16.0 | 73.0 |
| 4 | Plumbing leaks, minor | 60 | 10.0 | 83.0 |
| 5 | Pump seal failure | 48 | 8.0 | 91.0 |
| 6 | Control panel and sensor faults | 36 | 6.0 | 97.0 |
| 7 | Chiller compressor failure | 12 | 2.0 | 99.0 |
| 8 | MV switchgear fault | 6 | 1.0 | 100.0 |
| Total | 600 | 100.0 | ||
Read that chart and the conclusion writes itself: lighting is the problem, it is a third of all corrective work, go and fix lighting. The top three categories carry seventy-three percent of events, which looks reassuringly close to the familiar heuristic, and a planner under time pressure stops reading there.
Now take the exact same eight categories and the exact same events, and rank them by total unplanned downtime hours instead. Again, my own illustrative figures:
| Rank | Failure category (illustrative) | Events | Downtime hrs | % of hrs | Cumulative % |
|---|---|---|---|---|---|
| 1 | Chiller compressor failure | 12 | 456 | 38.0 | 38.0 |
| 2 | MV switchgear fault | 6 | 288 | 24.0 | 62.0 |
| 3 | Pump seal failure | 48 | 192 | 16.0 | 78.0 |
| 4 | Control panel and sensor faults | 36 | 108 | 9.0 | 87.0 |
| 5 | AHU filter and belt issues | 96 | 96 | 8.0 | 95.0 |
| 6 | Plumbing leaks, minor | 60 | 30 | 2.5 | 97.5 |
| 7 | Lighting and lamp failure | 210 | 21 | 1.8 | 99.3 |
| 8 | Door hardware and closers | 132 | 9 | 0.7 | 100.0 |
| Total | 600 | 1,200 | 100.0 | ||
The chart has inverted. Lighting, which was thirty-five percent of the work and the obvious priority, is seventh out of eight on availability impact and contributes under two percent of lost hours. Chiller compressor failure, which was seventh on the count chart at twelve events and would have been dismissed as rare, is now the single largest contributor at thirty-eight percent. MV switchgear, six events across the whole year, is second.
Both tables are correct. Both are the same underlying data. They simply answer different questions, and if you only ever run the first one you will spend the year optimising lamp replacement while the assets that actually take buildings offline stay exactly as they were. That contrast is the most important thing in this article, and it is available to anyone with two queries instead of one.
The useful synthesis is to read both together. Lighting and door hardware are a volume problem: they are eating technician capacity, and the fix is process, spares strategy or a shift to planned replacement, not reliability engineering. Chiller compressors and switchgear are an impact problem: low frequency, high consequence, and the fix is condition monitoring, criticality-driven PM and spares availability. Pump seals appear in the top five on both measures, which makes them the easiest case of all to argue for. Capturing the downtime that makes the second table possible is itself a discipline, covered in the backlog and downtime tracking guide.
6. Categorisation quality is the binding constraint
A Pareto chart is a picture of your coding scheme. Whether it is also a picture of your plant depends entirely on how disciplined that coding is. This is the constraint that binds before any of the others, and it is the one least often checked before presenting the chart.
The same failure patterns recur consistently in maintenance history:
- Default and catch-all codes. If the cause field is mandatory but the picklist is awkward, technicians pick the first option, or "other", or whichever code closes the ticket fastest. A chart with a large "other" bar is telling you about your data entry, not your equipment.
- Inconsistent granularity. One code says "electrical", the next says "contactor coil open circuit". They cannot be ranked against each other meaningfully, and the coarse code will always look bigger simply because it swallows more.
- Symptom recorded as cause. "Not working", "no cooling", "noise" are problem codes masquerading as cause codes. A Pareto chart of symptoms ranks complaints, which is a legitimate thing to rank but must not be presented as a ranking of causes.
- Drift over time. Codes get added, retired and reinterpreted. A twelve month window that spans a coding change contains two incompatible datasets stacked on top of each other.
- Missing data treated as zero. If half your work orders have no downtime recorded, a downtime-ranked Pareto is a ranking of the assets whose downtime somebody happened to log.
This is why I treat coding discipline as a prerequisite rather than a nice-to-have. The structure that makes this work is a separated Problem, Cause and Action scheme, with a constrained picklist per asset class, which is exactly what the failure codes guide sets out. Get that right and every subsequent analysis, Pareto included, becomes worth doing. Get it wrong and you are ranking artefacts. For the wider taxonomy question, ISO 14224:2016 is the reference standard for collecting and exchanging reliability and maintenance data, originally developed for the petroleum and gas sector but widely borrowed elsewhere for its equipment and failure-mode taxonomy. Note for the record that OREDA, often mentioned alongside it, is a proprietary members-only database rather than a standard.
A validation step worth building in
Before interpreting any Pareto chart, check three things: the share of records in "other" or blank, the share of records with the measure field missing, and whether the coding scheme changed during the window. If any of those is material, report the chart with that caveat attached or do not report it at all. A chart presented without those checks is an opinion wearing a bar graph.
7. Bad actor analysis, the main maintenance application
The dominant use of Pareto analysis in maintenance is bad actor identification: finding the small population of assets, or asset classes, that generate a disproportionate share of the unplanned work, the downtime or the cost. It is the standard first move when someone asks where a reliability improvement programme should start.
A workable sequence:
- Rank by asset or asset class first, not by cause. You want to know which equipment to look at before you care why it is failing. Use downtime or cost, not count.
- Normalise where the comparison is unfair. Twenty pumps will generate more absolute failures than two chillers. Failures per unit, or per running hour, often tells a truer story than raw totals, particularly across mixed asset populations.
- Cross-check against criticality. A bad actor that is not critical may simply be tolerable. A critical asset that is not a bad actor still deserves its PM regime. The intersection of high contribution and high criticality is where the programme starts, and that ranking comes from a separate exercise, set out in the equipment criticality analysis guide.
- Pick a defensible cut-off. Wherever the cumulative curve flattens, plus a sanity check that the list is short enough to actually work on. Three to five bad actors is a programme. Twenty is a wish list.
- Hand each one to a causal method. Pareto has now done its job. It has told you where to look.
Re-run the analysis on the same measure, same categorisation, same window length, at a regular interval. The chart moving is your evidence that the programme worked, and it is far more persuasive than a percentage improvement claim, because the audience can see the bar that used to be first and is now fourth. Pair it with schedule compliance and the rest of the preventive maintenance KPI set so that the improvement story has more than one line of evidence behind it.
8. Using it iteratively: a Pareto within a Pareto
A single Pareto chart usually lands you on a category too broad to act on. "HVAC" is not an improvement project. The standard move, and the one that makes the technique genuinely powerful, is to take the top category and run the same analysis inside it.
A worked drill-down, using the illustrative downtime table above as the starting point:
↓ top bar = chiller compressor failure, 456 hrs (illustrative)
Pareto 2: chiller compressor events by individual chiller unit
↓ top bar = two units out of eight carry most of the hours
Pareto 3: those two units, by recorded cause code
↓ top bar = one dominant cause, e.g. refrigerant loss
↓
Now switch tools: 5 Whys or fishbone on that specific cause
Three iterations took you from twelve hundred hours of undifferentiated downtime to one cause on two named units, which is small enough for a causal investigation to actually succeed on. Two cautions on drilling. First, each level cuts your sample, so by the third iteration you may be ranking single-digit event counts, where the ordering is not meaningful and you should stop treating the ranking as evidence. Second, resist drilling down the same branch out of momentum: at each level, re-ask whether the second bar is now close enough to the first that it deserves attention too.
9. Where Pareto analysis is the wrong tool
Every technique has a domain, and this one is used well outside its own more often than most. The cases where I would not use it:
- Too few categories. With three or four categories there is nothing to prioritise. Sorting four bars adds no information you did not have from reading the four numbers, and dressing it up as a Pareto chart implies an analytical rigour that is not there.
- Genuinely even distribution. If the cumulative curve is close to a straight diagonal, there is no dominant minority and no cut-off to find. That is a real and reportable result, and the honest conclusion is that improvement here requires a systemic change rather than targeted attack on a few categories.
- Small samples. With a handful of events per category, the ranking is mostly noise, and re-running it next quarter will produce a different order for no substantive reason.
- Anything needing causation. Pareto ranks correlation with a total. It says nothing about mechanism. Acting on the top bar without understanding why it is the top bar is how organisations fix symptoms repeatedly.
- Trend questions. A Pareto chart is a snapshot of a window. It cannot show you that something is deteriorating. Use a time series for that and a Pareto for allocation.
Never use a count-based Pareto to deprioritise a safety risk
This one is firm, and it is the most serious misuse of the technique. Low-frequency high-consequence events, a gas release, an arc flash, a fire pump failing to start, a lift entrapment, will always sit in the tail of a frequency chart. They occur rarely by definition. A Pareto ranked by count will therefore always rank them as unimportant, and that ranking is an artefact of the measure, not a statement about risk. Safety and life-safety decisions are made on consequence and likelihood together, through risk assessment, not on where a bar falls in a frequency chart. IEC 31010:2019 "Risk management, risk assessment techniques" catalogues the methods suited to that question. If you present a Pareto chart to a safety committee, say explicitly that it is not a risk ranking.
10. How it combines with the other tools
Pareto analysis is a triage tool and it works best chained to something that explains mechanism. The pairing is so standard that the two are usually taught together: Pareto to choose where to look, then a causal method to understand what you are looking at.
On standards: IEC 62740:2015 "Root cause analysis (RCA)" is a real international standard for after-the-event causal analysis. It sets out principles and process steps and describes a set of named recognised techniques, including the "Why" method and the Ishikawa or fishbone diagram, so it is simply wrong to say RCA has no standard. As far as I am aware no standard prescribes how to run a Pareto analysis specifically, which is unsurprising for what is essentially a sort and a running total, and I would not make a stronger claim than that either way. What matters practically is the handover between tools:
- Pareto then fishbone. Once Pareto names the category, a cause-and-effect diagram structures a group's thinking across the plausible cause families without committing to one too early. This is the most natural partner, and the fishbone diagram guide covers the mechanics.
- Pareto then 5 Whys. For a single clear failure with a traceable chain, iterative questioning is faster and lighter than a full workshop. See the 5 Whys guide, including its known weakness of following one chain and missing parallel causes.
- Pareto then FMEA. Where the top category is an asset class rather than an event, a failure modes and effects analysis on that class is the systematic follow-through.
- Pareto after the fix, as verification. Re-running the same chart after the intervention is the cheapest proof of effect available.
- Pareto alongside criticality. Contribution and consequence are different axes. Plot both and the genuine priorities separate themselves out.
For the wider set and how they fit together, the RCA tools overview maps the options against the kinds of problem each suits.
11. Where the data comes from, practically
None of this works without a data source, and for a maintenance team that source is the work order history in the maintenance system of record. A Pareto analysis is, in query terms, a group-by on a cause or asset field with a sum over a measure field, filtered to a date range and a work type. Any competent maintenance system can produce that, either through its own reporting layer or by pushing the work order table into a spreadsheet or BI tool.
Two practical observations. First, the constraint is rarely the reporting tool, it is whether the cause code and downtime fields were populated consistently at the point of work order closure, which is a process and configuration question rather than an analytics one. Second, bear in mind that a built-in "top failure causes" widget has already silently chosen a measure on your behalf, and in practice it has usually chosen count, which is the measure most likely to mislead. Find out what it is summing before you quote it. If you are still deciding what a maintenance system needs to do for you, the CMMS introduction covers the data capture side of the requirement.
The idea to walk away with
Pareto analysis is a sorting tool wearing the costume of an insight. Its value is entirely in forcing a ranked, evidence-based conversation about where to start, and its danger is entirely in how easily the ranking is mistaken for a verdict. The real skill is not drawing the chart, which takes minutes; it is choosing the measure deliberately, knowing what that measure hides, checking that the categorisation underneath it means anything, and handing the winning category to a tool that can actually explain it.
If you take one habit from this article, take the double run: every Pareto analysis on two measures, side by side, before anyone acts on it. In the illustrative example above that one extra query moved the priority from lamp replacement to chiller compressors. That gap between the two answers, both drawn from the same records, is the whole argument.
Final thoughts
The 80/20 framing has done the technique a disservice by making it sound like a finding rather than a method. Drop the number, keep the method. Compute your own split, report it honestly, state your measure, and be willing to report that your data is evenly distributed and Pareto is not the right tool here. That kind of reporting builds the credibility that makes the next recommendation land.
And hold the safety line firmly. A frequency-ranked chart will always rank rare catastrophic events last, and no amount of analytical enthusiasm makes that a reason to deprioritise them. Pareto is for allocating improvement effort across chronic problems. Consequence-based risk assessment is for everything that can hurt someone. Keep those two jobs separate and the technique will serve you well for as long as you use it.
Disclosure
Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.
Not sure your failure data can carry an analysis yet?
Independent advisory on failure coding structure, downtime capture, bad actor analysis and the reporting layer that makes reliability improvement measurable. 22+ years across utilities, oil and gas, manufacturing, government and facility operations.
Book a conversationRelated reading: Root cause analysis: methods and step-by-step guide, Fishbone diagram for root cause analysis, 5 Whys: method and examples, RCA tools: 8 methods explained, Failure codes: Problem, Cause, Action, Maintenance backlog and downtime tracking, Equipment criticality analysis.
Muhammad Abbas
CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.
Work with me