Most asset-intensive organisations I have worked with own a criticality matrix. Far fewer can explain why theirs has five bands rather than four, why consequence is weighted the way it is, or what a score of twelve actually means next to a score of ten. The matrix arrived with a consultant, got populated in a fortnight, and has sat in a spreadsheet ever since, quietly driving preventive maintenance frequencies and spares holdings that nobody revisits. This article is about the instrument itself: how to build one, how to populate it, how to read it, and where its mathematics does not deserve the confidence people place in it.
The message up front: a criticality matrix is a sorting and conversation tool, not a calculator. It is very good at separating the equipment that clearly matters from the equipment that clearly does not, and at forcing a structured argument in a room full of engineers. It is poor at telling you that asset A at fifteen points is meaningfully riskier than asset B at fourteen. Build it for the first job, design the scale and the override rules yourself, and never let the arithmetic outrank engineering judgement.
1. What a criticality matrix is, and the shape it usually takes
A criticality matrix is a grid that positions each item of equipment according to two axes, and reads a criticality band off the cell where the item lands. In maintenance and asset management the two axes are typically some form of consequence and some form of likelihood.
- The consequence axis answers: if this asset stops performing its function, how bad is that, in terms of safety, statutory compliance, service loss, environmental release, cost and reputation?
- The likelihood axis answers: how often is that functional loss expected? Most organisations cannot state a genuine probability, so they use a proxy: historical failure frequency, a condition or age rating, a duty-severity rating, or a judgement band from "rarely" to "several times a year".
- The cell where the two meet maps to a criticality band, and the band drives maintenance strategy, spares policy, response priority and monitoring investment.
Some organisations run a single-axis matrix that scores consequence only. That is a legitimate simplification when your failure history is too poor to support any likelihood estimate, and it is more honest than inventing frequencies. It does mean you lose the ability to distinguish "protect this" from "fix this". If you go single-axis, say so explicitly.
The matrix is one of the named techniques in the international catalogue of risk assessment methods, IEC 31010:2019 , "Risk management - Risk assessment techniques". Note the designation: it is IEC 31010, not ISO 31010 and not ISO/IEC 31010, because only the withdrawn 2009 edition carried the dual prefix. For the wider governance frame, ISO 31000:2018 "Risk management - Guidelines" is the reference, but be clear with management that it is guidance with no auditable requirements, so there is no accredited ISO 31000 certification for an organisation.
This article is deliberately about the instrument. The method around it, the consequence dimensions in depth, who sits in the workshop and how the exercise is governed all belong to the equipment criticality analysis complete guide. Read that for the exercise, this for the grid.
2. Two different matrices that get confused constantly
There are two grids in circulation that look almost identical and do entirely different jobs, and I have sat in meetings where the same spreadsheet was being used for both.
- An equipment criticality matrix ranks assets. The unit of analysis is a pump, a chiller, a switchboard, a lift. The output is a standing classification on the asset record that drives maintenance strategy for years, reviewed annually or on change.
- A risk assessment matrix ranks activities and hazards. The unit is a task: working at height on that chiller, isolating that switchboard, entering that tank. The output is a control decision for a specific job, recorded on a method statement or permit and reassessed whenever conditions change.
The confusion matters because the two have different owners and different review cycles. A maintenance planner owns the asset criticality matrix; a safety function owns the task risk matrix. If you want the health and safety side of the family, that is covered in how to calculate risk with a risk assessment matrix. Everything below is about ranking equipment.
3. Choosing the axes, and the cost-on-an-axis mistake
A common and damaging design error is putting asset replacement cost or capital value on the consequence axis. It is understandable: the value is already in the fixed asset register, it is objective, and it needs no workshop to establish. It is also the wrong variable.
Criticality is about the consequence of losing the asset's function, not what the asset cost to buy. An inexpensive differential pressure switch that trips a whole water treatment train on failure is more critical than a costly standby transformer with a fully rated redundant partner. Rank by capital value and you invert both. The matrix then sends your preventive maintenance effort, spares budget and condition monitoring at the expensive equipment rather than the consequential equipment, which is precisely the failure the exercise exists to prevent.
The test for a consequence axis
Ask of every candidate dimension: does this describe what happens to people, compliance, output or the environment when the asset stops doing its job? If it describes a property of the asset itself, such as its price, age, manufacturer or size, it does not belong on the consequence axis. Age and condition may legitimately belong on the likelihood axis. Price belongs nowhere in criticality, only in the later business case.
On the likelihood axis, be honest about what you have. Genuine probability requires a population of similar assets, a clean failure history and consistent failure coding, and most organisations have none of the three. What goes on the axis is a proxy, and it should be labelled as one. Historical failure counts per asset class are the best available proxy where the data is trustworthy; where it is not, a banded engineering judgement of expected frequency is defensible so long as the bands are defined in words that mean the same thing to everyone in the room.
One more axis decision worth making explicitly: whether redundancy is handled on the consequence axis or as a separate modifier. Either treatment works. What does not work is handling it inconsistently, so that some assessors discount for redundancy and others do not. Write the rule down.
4. Designing the scale: how many bands, and odd against even
There is no correct number of bands, and I am not going to give you one. What I can give you is the trade-off, so you can choose deliberately rather than inherit somebody else's choice.
- Fewer bands are faster to populate and more robust: two assessors are more likely to agree on "high or low" than on "four or five out of seven". The cost is resolution. A matrix with only three bands and half the register in the middle one has not told you much.
- More bands give finer discrimination but invite false precision. Once you have seven or nine, assessors argue about single-point differences the evidence cannot support, and the workshop slows to a crawl.
- An even number removes the safe middle. With four or six bands there is no neutral option, so an assessor who is unsure has to commit to the upper or lower half. That forces the argument that produces the useful information. It is genuinely harder work.
- An odd number is easier to run in a workshop. People are comfortable with a middle, the session moves faster, and an honest "medium" is sometimes correct. The price is central tendency: a large share of assets drifting into the middle band because it is the least contestable place to put them.
Pick the smallest number of bands that still separates the decisions you will actually make, working backwards from the decisions rather than forwards from the scale. Then write anchored definitions for every band on every dimension, in operational language, before anyone scores anything. "Consequence band 3: loss of service to a single building for under four hours, no injury, no reportable release, recoverable within the shift." That is scoreable. "Consequence band 3: moderate" is not, and a matrix built on unanchored adjectives will not be reproducible between assessors or between years.
5. The ordinal problem: why the arithmetic does not really work
This is the honest core of the article, and the part most criticality training leaves out. Consequence and likelihood scores are ordinal: they tell you the order of things, not the distance between them. Band 4 is worse than band 3, and band 3 is worse than band 2, but nothing in the design says the gap between 3 and 4 is the same size as the gap between 2 and 3. In most real scales it is not, because consequence bands tend to be defined across wildly different magnitudes, so moving up one band near the top can represent many times the impact of moving up one band near the bottom.
The implication is uncomfortable. Multiplying two ordinal ranks produces a number with no genuine arithmetic meaning. It is not a risk quantity, and it cannot be averaged, compared as a ratio, or summed across a portfolio in any defensible way. Two assets scoring twelve can represent very different risks: consequence 3 times likelihood 4 is a nuisance that happens often, consequence 4 times likelihood 3 may be a safety event that happens occasionally, and treating them as equivalent because the product matches is exactly the error the matrix invites.
The same defect as the RPN in FMEA
If this sounds familiar, it should. It is precisely the criticism levelled at the Risk Priority Number in failure mode and effects analysis, where severity, occurrence and detection ranks are multiplied into a single figure. The mathematics is no sounder there, which is why the current international standard for the technique, IEC 60812:2018 "Failure modes and effects analysis (FMEA and FMECA)", and the industry handbooks that followed it, have moved away from ranking purely by RPN toward structured action-priority logic. Readers should see the connection, because the fix is the same in both cases: use the score to sort and to start the argument, then apply explicit rules about what must be escalated regardless of the arithmetic. The FMEA guide covers that side in detail.
None of this makes the matrix useless. It makes it a different tool from the one people think they are holding. A criticality matrix is excellent at three things: separating the clearly critical from the clearly trivial; surfacing disagreement between operations, maintenance and engineering about what matters, often the most valuable output of the exercise; and producing an auditable record of the reasoning behind a maintenance decision. It is bad at one thing, fine-grained ranking within a band, so do not use it for that. Group into bands, act on the bands, and use engineering judgement inside them.
I would also stop presenting the score as a measured quantity. Report the band and keep the component scores visible alongside it, because high consequence with low frequency calls for protection and spares, while low consequence with high frequency calls for root cause work. A single product hides that distinction.
6. Weighting multiple consequence dimensions
Consequence is rarely one thing. A realistic matrix scores several dimensions: commonly safety, statutory exposure, operational loss, environmental impact, cost, and sometimes customer impact. That raises the question of how to combine them, and there are three sensible answers.
- Take the highest. The consequence score is the worst of the dimension scores. Simple, transparent and conservative, and it never lets a severe safety consequence be diluted by a trivial cost consequence. The drawback is discarded information: an asset severe on one dimension looks identical to one severe on all five.
- Weighted sum or average. Assign each dimension a weight, multiply, add. This retains more information and lets the organisation express what it cares about. It is also the route through which severe consequences get averaged into invisibility, which is why it needs the override rules in the next section.
- Highest, with a weighted tie-break. Use the worst dimension to set the band, and a weighted sum only to order assets within a band for scheduling. This is what I usually recommend: it keeps the conservatism where it matters and puts the softer arithmetic where being slightly wrong costs little.
On the weights themselves: I will not publish a set, and you should be suspicious of anyone who does. Weights are not a technical parameter. They are a statement of organisational values, and properly the business's decision rather than the reliability engineer's. Deciding that operational loss outweighs environmental impact is a value judgement with real consequences, and it deserves to be made consciously by people with the authority to make it.
The practical implication is procedural. Weights should be written on a single visible page, with a sentence of rationale each, dated, and signed off by an accountable owner: an asset management lead, an operations director, a risk committee. What they must not be is a row of hidden multipliers in a locked spreadsheet column that nobody remembers agreeing. A weight set that cannot be traced to a named person will be quietly retuned the first time someone dislikes an answer.
7. Override rules: do not let averaging hide a safety consequence
Some consequences should not be negotiable. If the failure of an asset can injure somebody, breach a statutory duty, or defeat a life safety system, that asset needs to sit in a high band regardless of how rarely the failure is expected. Likelihood is the wrong lever to pull against an intolerable consequence, and multiplication pulls it anyway: a top consequence rank times the lowest likelihood rank produces a modest product, which is how a fire pump or an emergency generator ends up scored below a production conveyor.
The fix is not to fiddle with the weights until the answer looks right. It is an explicit override layer sitting on top of the arithmetic, documented as part of the matrix design. Overrides I would build into almost any equipment criticality matrix:
- Safety escalation. Any asset whose functional failure can credibly cause injury escalates into the top bands regardless of the calculated product.
- Statutory and life safety escalation. Any asset subject to mandated inspection or testing, or whose failure breaches a permit, licence or code obligation enforced by your authority having jurisdiction, escalates on the same basis. Those duties differ by jurisdiction, so this list has to be drawn up locally rather than copied from another country's matrix.
- Environmental release escalation. Any asset whose failure can cause a reportable release escalates.
- Single point of failure escalation. Any asset with no redundancy, no bypass and no realistic workaround escalates by one band, because the consequence of its loss is not recoverable by operational means.
Record overrides as overrides on the asset record, not baked invisibly into an adjusted score. When someone reviews the matrix in three years, the difference between "this scored high" and "this was escalated because it is a statutory life safety asset" is the difference between a decision they can audit and a number they have to take on trust.
Where overrides stop being useful
There is a limit. If your override rules escalate a large share of the register into the top band, the matrix has stopped discriminating and you have simply moved the problem. That usually means the consequence bands are too coarse at the severe end, or the override criteria are drawn too broadly, for example escalating every asset that appears anywhere in a statutory schedule rather than those whose failure actually breaches a duty. Tighten the criteria before you widen the bands.
8. A worked example matrix
What follows is invented for illustration. The scale, bands, implied weighting and assets are all hypothetical, chosen to show the mechanics. They are not a recommended scale and not benchmarks: design your own against your own operation, service obligations and jurisdiction, and the result will not look like this. Here I use a hypothetical four-point consequence scale (C1 lowest to C4 highest), a hypothetical four-point frequency proxy (F1 rarely to F4 often), a "take the highest dimension" consequence rule, and the override rules above. Four bands appear only because an even scale removes the safe middle, not because four is correct.
| Illustrative asset | Worst consequence dimension | C (illustrative) | F (illustrative) | Raw product | Override applied | Final band |
|---|---|---|---|---|---|---|
| Hypothetical electric fire pump, no standby | Statutory / life safety | C4 | F1 | 4 | Statutory and single point of failure | Band A (highest) |
| Hypothetical sole process chiller serving a data hall | Operational loss | C4 | F2 | 8 | Single point of failure | Band A |
| Hypothetical effluent transfer pump, duty of two | Environmental release | C3 | F3 | 9 | Environmental release | Band A |
| Hypothetical packaging line conveyor | Operational loss | C3 | F4 | 12 | None | Band B |
| Hypothetical standby generator with rated twin | Statutory / life safety | C4 | F1 | 4 | Statutory (no single point escalation) | Band B |
| Hypothetical air handling unit, one of twelve on a floor | Operational loss | C2 | F3 | 6 | None | Band C |
| Hypothetical office lighting circuit, local | Operational loss | C1 | F2 | 2 | None | Band D (lowest) |
| Hypothetical car park extract fan, non-ventilation-critical | Operational loss | C1 | F3 | 3 | None | Band D |
Read down the raw product column and the ordinal problem is visible at a glance. The hypothetical fire pump and standby generator both score 4, the lowest products in the table. The hypothetical conveyor scores 12, the highest, on a purely operational consequence recoverable within a shift. Sort by product and work down the list and you would attend to the conveyor first and the fire pump last. The override layer prevents that, which is why the override column sits on the same page as the score rather than in a separate procedure nobody opens. Notice too the two C4 statutory assets landing in different bands purely because one has a rated twin: that is the redundancy rule working, and it is defensible only because the rule was written down before scoring started.
9. The design choices, and what goes wrong if you get them wrong
Building a matrix is a sequence of design decisions, each with its own failure mode. This is the checklist I would work through before a single asset is scored.
| Design choice | Options | Consequence of getting it wrong |
|---|---|---|
| Consequence axis basis | Consequence of functional loss, or asset replacement cost | Cost-based ranking inverts the register, sending maintenance effort and spares at expensive equipment instead of consequential equipment |
| Second axis | True probability, historical failure frequency, condition or age proxy, duty severity, or no second axis at all | An invented probability looks rigorous and is not; no second axis means you cannot tell "protect this" from "fix this" |
| Number of bands | Fewer for robustness, more for resolution | Too few and everything clusters with no actionable distinction; too many and assessors argue over differences the evidence cannot support |
| Odd or even band count | Even removes the safe middle, odd is easier to run | Odd scales attract central tendency, with a large share of assets parked in the middle; even scales slow the workshop and irritate assessors |
| Band definitions | Anchored operational descriptions, or adjectives such as low, medium, high | Unanchored adjectives are not reproducible between assessors or between years, so the matrix cannot be audited or repeated |
| Combining consequence dimensions | Take the highest, weighted sum, or highest with a weighted tie-break | Averaging dilutes a severe single-dimension consequence into a comfortable middle band |
| Weight ownership | Published and signed off by an accountable owner, or embedded in a spreadsheet | Hidden weights get retuned until the output matches somebody's prior opinion, and nobody can show they did not |
| Override rules | Explicit escalation for safety, statutory, environmental and single point of failure, or none | Without overrides, rare but intolerable consequences score low and rank below frequent nuisances |
| Redundancy treatment | Discount on the consequence axis, or a separate documented modifier | Inconsistent treatment between assessors makes scores incomparable across the register |
| Unit of assessment | Functional location, equipment item, or assembly | Mixed granularity produces a register where a whole chiller plant and a single valve carry the same weight |
| Review trigger | Fixed cycle plus change-driven review, or no trigger | A matrix that is never revisited drives PM frequencies and spares on an operating reality that no longer exists |
10. Populating it without the exercise stalling
Matrix design is a day or two of careful thought. Population is where criticality exercises die, usually when workshop fatigue sets in and the remaining register gets scored by one tired person on a Friday afternoon. The sequence I have seen hold up:
- Agree calibration examples first. Take a handful of assets everybody already has an instinct about and score them together. The purpose is not the scores, it is to discover that two assessors read "moderate operational loss" differently while that is still cheap to resolve. Write the agreed examples into the band definitions as anchors.
- Pilot a small representative set. A few dozen assets spanning your classes and expected bands. Score them, then look at the distribution and the disagreements. This is where you find out the scale has no discrimination at the bottom, or the redundancy rule is ambiguous. Fix the instrument now, because after the bulk run a design change means rescoring everything.
- Then run the bulk by asset class, not by location. Scoring all the pumps together, then all the air handling units, is faster and more consistent, because the assessor stays in one frame and inherits the previous decision as a reference. Building-by-building scoring produces drift.
- Template the repeats. Where many near-identical assets sit in similar service, score the class once, apply it as a default, and flag only the exceptions, which are usually about service context rather than the equipment.
- Keep a decisions log. Every time the room resolves an ambiguity, write the ruling down in one line. That log is what lets a different assessor reach the same answer next year.
- Timebox and accept provisional scores. An asset scored approximately today and corrected at the first review beats one left blank while the perfect scale is debated. Mark provisional scores as provisional so the correction happens.
The two things that reliably stall the exercise are trying to perfect the scale during the bulk run, and letting the meeting relitigate the band definitions on every contentious asset. Doing the calibration and the pilot properly prevents both.
11. Reading the output into decisions
A matrix that changes nothing is a filing exercise. The output should drive four distinct decision families, and the bands are only defensible if each leads somewhere different.
- Maintenance strategy. High bands justify condition-based or predictive approaches and a full failure-mode analysis. Middle bands usually land on preventive maintenance at a defined interval. Low bands can often run to failure with a fast corrective response, which is a legitimate strategy rather than neglect. The selection logic belongs to reliability centred maintenance, and the criteria defining what may legitimately be called RCM are set out in SAE JA1011_202411. See the RCM introduction and the complete guide to preventive maintenance.
- Spares and stocking policy. Criticality is what justifies holding an expensive part that may never be used. High bands with long lead times are the classic critical spare; low bands with short lead times should not be stocked. See spare parts and MRO inventory in a CMMS.
- Response priority and target times. Criticality band combined with fault severity is the standard basis for work order priority and rectification targets, whether those sit in an internal service standard or a contracted service level agreement.
- Monitoring investment. Spend on sensors, condition monitoring and analytics should track the top bands, not the whole register. Criticality is the filter that stops a monitoring programme spreading thin across equipment whose failure nobody would notice.
A useful discipline is to require that every band maps to a named, written treatment. If two adjacent bands lead to the same maintenance strategy, stocking rule, response target and monitoring decision, they are not two bands. Merge them. Most matrices are one or two bands more complex than the decisions they inform.
For the wider engineering frame around all of this see the reliability engineering complete guide; for the failure-mode vocabulary that feeds the consequence assessment, common failure modes and how to analyse them; and for a short reference on how the resulting bands are described, asset criticality classification.
12. How criticality matrices fail in practice
The instrument is sound enough. The failures are almost all organisational, and consistent enough that you can check for them directly.
- Score inflation. Nobody wants to be the person who rated an asset low, so scores drift upward over successive reviews. Once everything is important, nothing is prioritised. The tell is a distribution that has migrated toward the top with no change in the operation. The defence is anchored band definitions with reference examples, and a reviewer who must justify every upward movement.
- Weights tuned until the answer looks right. The output contradicts somebody's prior conviction, so the weights get adjusted until the matrix agrees. This is the most corrosive failure, because it destroys the value of the exercise while leaving it looking rigorous. Published, dated, signed-off weights are the only real protection, with a rule that weight changes trigger a full rescore and a visible comparison against the previous run.
- Everything in one band. A matrix where most of the register sits in a single band has not discriminated. The cause is usually unanchored adjectives, an odd scale attracting central tendency, or override criteria capturing most of the estate. That is a scale design problem, not a scoring problem.
- The matrix nobody revisits. Assets get added, redundancy gets removed during a project, a building changes use, and the matrix still reflects the operation as it was on the day of the original workshop, quietly driving PM frequencies and spares against a reality that has moved. It needs a review trigger on change, not only a calendar cycle, and new assets should be scored at commissioning.
- Criticality that never reaches the maintenance system. A band that lives in a consultant's spreadsheet and is never written onto the asset record changes nothing. Whatever platform you run, the criticality field should be populated, reportable and referenced by the priority and planning logic. Any competent system supports this; the gap is process, not software.
- Treating the score as a measurement. League tables ordered by product, portfolio averages, or comparisons between organisations with different scales all read arithmetic meaning into ordinal ranks that do not carry it.
The idea to walk away with
A criticality matrix earns its place by forcing a structured argument and producing a coarse, defensible sort. It does not earn its place as a calculator. The scores inside it are ordinal ranks, the products of those ranks are not risk quantities, and the moment you treat single-point differences as meaningful you are reading precision the instrument never had.
So build it deliberately: consequence of functional loss on the consequence axis and nothing else, the smallest band count that separates decisions you will actually take, anchored operational band definitions, weights that are a visible signed-off value judgement rather than a hidden spreadsheet column, override rules so safety and statutory consequences escalate regardless of likelihood, and a review trigger so it reflects the operation you have rather than the one you had.
Final thoughts
The most valuable hour in a criticality exercise is usually the calibration session, where three people who all thought they agreed discover they do not. That conversation is the real product. The grid is just the structure that makes it happen in a consistent order and leaves a record of what was decided.
If your organisation already has a matrix, the audit is short. Are the band definitions operational rather than adjectival? Who signed off the weights, and when? Are the override rules written down and visible on the asset records they escalated? Has the distribution moved since the original run, and can anybody explain why? Was the last review a calendar event or a response to a change in the operation? A matrix that survives those five questions is doing real work. One that does not is a spreadsheet with authority it has not earned.
Disclosure
Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.
Designing or auditing a criticality matrix?
Independent advisory on criticality scale design, weighting governance, override rules, population workflow and getting the resulting bands to actually drive maintenance strategy, spares and response priority in your maintenance system. 22+ years across utilities, oil and gas, manufacturing, government and facility operations.
Book a conversationRelated reading: Equipment criticality analysis: complete guide, Risk assessment matrix: how to calculate risk, FMEA: failure mode and effects analysis guide, Common failure modes and how to analyse them, What is reliability engineering, Asset criticality classification, RCM introduction, Spare parts and MRO inventory in a CMMS, Preventive maintenance: the complete guide.
Muhammad Abbas
CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.
Work with me