Most organisations run a risk matrix, and few can say where their particular one came from. It was inherited, copied from a template, or drawn on a whiteboard by somebody who has since left. The bands have no written descriptors, the scores carry no reasoning, and two competent supervisors rating the same task will land in different cells and both be confident. None of that makes the matrix useless. It means it is being asked to do a job it was never able to do. A risk assessment matrix is a grid combining a judgement about likelihood with a judgement about severity to produce a rating that drives a decision.
The message up front: a risk matrix is a communication and prioritisation device, not a measurement instrument. It records two human judgements in a form a group can argue about productively and a decision-maker can act on. Treat the output as a structured opinion with a documented rationale and it is one of the most valuable tools in safety practice. Treat it as a calculation that produces a fact and it will mislead you in predictable ways.
One qualification, stated once. The scheme your organisation uses, the bands, the descriptors, the boundaries and the actions each band triggers, should be set by a competent safety professional against the law applicable in your jurisdiction and your organisation's own stated risk appetite. This article explains how the instrument works and where it fails. It is not a scheme you can adopt.
The scope is narrow. The end-to-end process is covered in the complete risk assessment guide; controls in the hierarchy of controls guide; and if hazard and risk are still doing double duty in your paperwork, start with hazard versus risk. What follows is only about the matrix: how it is built, how a rating is arrived at, and how it deceives.
1. What a risk matrix is, and what it is for
Strip away the colouring and a risk matrix is a two-dimensional lookup table. One axis carries ordered bands of severity, the other ordered bands of likelihood, and each cell is assigned a rating: a level, a colour, a score, or all three. You judge each axis, read off the cell, and the cell tells you what happens next.
The genuine value is social and procedural rather than analytical. A matrix forces a group to separate two questions they would otherwise blur together, to state a position on each in words that can be challenged, and to record a shared conclusion. It makes risk comparable enough that a manager with forty assessments can decide which three to deal with this week, and it leaves a record an auditor or investigator can read years later. None of that depends on the number being accurate.
What a matrix is not is a measurement. No instrument is applied; there is no calibration, no units, no error bar. The recognised international guidance on risk management, ISO 31000:2018, sets out principles and a framework but is explicitly guidance rather than auditable requirements, which is why there is no accredited ISO 31000 certification for an organisation and anyone offering you one is selling something else. The companion document for technique selection is IEC 31010:2019, "Risk management - Risk assessment techniques", which catalogues assessment methods and where each is appropriate. Note the prefix: it is IEC 31010, not ISO 31010. Only the withdrawn 2009 edition was dual-prefixed, and the wrong form appears constantly in HSE copy. Both are voluntary and neither is law anywhere by itself.
The test I apply
Remove the rating and leave only the hazard description, the controls and the reasoning. Is the document still useful? If yes, the matrix is doing its proper job as a summary. If the rating was the only substance, it has replaced the assessment rather than summarising it.
2. The severity axis: what exactly is being rated
Severity is the easier axis and even it is routinely muddled. In an occupational health and safety context the consequence being rated is harm to people, with bands running from a minor injury needing no treatment up to a fatality or multiple fatalities.
The complication is that organisations want one matrix to serve several purposes, so environmental consequence, asset damage, business interruption, regulatory exposure and reputational damage get folded onto the same axis, each band carrying parallel descriptors and the assessor taking the highest that applies. Its cost is rarely acknowledged: once different consequence types share one axis, two ratings in the same band are no longer comparable in any meaningful sense. A rating driven by potential permanent disability and one driven by a week of production loss sit in the same row and attract the same procedural response, which is defensible for prioritising spend and indefensible for anything touching human harm. Where I see this go wrong most often is a register in which safety and commercial risks were merged for reporting convenience, and the safety items have quietly been out-prioritised by larger financial numbers.
If you must run a combined axis, keep the consequence type that drove each rating visible in the record. One narrower point saves a lot of argument: rate severity on the credible worst outcome of the hazard as the task is actually performed, not the most likely outcome and not the theoretical maximum. A fall from a two-metre platform can credibly kill somebody, and rating it as a sprain because sprains are what usually happens is the commonest way severity gets understated.
3. The likelihood axis, and the three questions people confuse
Likelihood is the genuinely difficult axis, and this is the most useful thing on the page. Ask a working group how likely something is and they answer one of three different questions, often without noticing which, and the three produce materially different ratings. If your scheme does not say which question the band refers to, your assessments are not comparable with each other.
| The question actually being answered | What it measures | Why it is not the others |
|---|---|---|
| How likely is the hazardous event? | Probability the initiating event happens at all: the load slips, the guard is defeated, the valve is opened under pressure. | Says nothing about whether anybody is hurt. A load can slip into an empty exclusion zone. |
| How likely is harm, given the event? | Conditional probability of injury once it happens: presence, position, reaction time, protection in place. | Where most control effectiveness shows up. Rating this alone ignores how often the event occurs. |
| How often is the task done, or the person exposed? | Frequency and duration: once a year, weekly, hourly, continuously through a shift. | Exposure, not probability. It multiplies the other two rather than substituting for either. |
An illustration with my own clearly hypothetical judgements. Manual replacement of a filter in a pressurised line: the hazardous event, residual pressure released while breaking the joint, might reasonably be judged unlikely because an isolation and depressurisation procedure is in force; the conditional likelihood of harm given that release, with a technician at the joint and no face protection specified, might reasonably be judged very high; and the task is performed weekly. Three questions, three bands, and a group that has not agreed which one it is answering will produce a rating anywhere across half the matrix.
The fix costs nothing: state in the scheme which question the band refers to, and state it on the form rather than only in a procedure nobody reads. For task and activity risk the most useful convention I know is to rate the likelihood of harm occurring, the combined probability of the event happening and of it harming somebody, handling exposure explicitly rather than smuggling it in.
On task risk these judgements are rarely derived from failure data. They come from operational experience, near-miss history and the assessor's knowledge of how the job is really done rather than how the method statement says it is done. That is legitimate, and it is also why two teams diverge. Techniques that put more rigour behind a likelihood judgement, fault trees, event trees, bowtie diagrams, layers-of-protection analysis, are catalogued in IEC 31010:2019; a matrix is one of the simplest entries in that catalogue. If a likelihood judgement is carrying a very high-consequence decision, the matrix is the wrong instrument.
4. Exposure: the hidden third dimension
A two-axis matrix has a structural gap. Take two tasks with identical severity and identical conditional likelihood of harm, one performed once a year during a shutdown and one every hour of every shift. They do not carry the same risk to the workforce, and a plain likelihood-by-severity matrix places them in the same cell.
This is why some schemes add a third axis, variously called exposure, frequency or probability of exposure. Machinery risk assessment has handled it explicitly for a long time: ISO 12100:2010, "Safety of machinery - General principles for design - Risk assessment and risk reduction", which consolidated and replaced three earlier documents and is currently under revision, treats the probability of occurrence of harm as a composite including frequency and duration of exposure alongside the probability of a hazardous event and the possibility of avoiding or limiting harm. That decomposition is useful well outside a machinery context, because it names what a two-axis matrix silently merges.
Adding an axis is not free: more cells, more boundary arguments, and a heavier form that supervisors start filling in mechanically. If the work you assess ranges from annual shutdown tasks to continuous operations, exposure deserves to be explicit. If almost everything is a routine daily task, folding it into the likelihood descriptors and saying so in writing is defensible. What is not defensible is having no position at all, which is where most inherited matrices sit.
5. Scale design, and why there is no authoritative scale
The 5x5 risk matrix is so widespread that people assume it is specified somewhere. It is not. The number of bands, the descriptors, the boundaries between them and the rating assigned to each cell are local design choices. No standard I am aware of prescribes a particular matrix, and the guidance that discusses matrices describes them as a technique rather than publishing one for adoption. So you cannot copy a matrix out of an article, including this one, and treat it as compliant with anything. That is why there is no recommended scale on this page. The choices you have to make:
| Design choice | Options | Consequence of getting it wrong |
|---|---|---|
| Bands per axis | 3x3, 4x4, 5x5 or more | Too few and everything clusters mid-scale. Too many and assessors cannot tell adjacent bands apart, so the resolution is noise presented as precision. |
| Even or odd band count | Odd gives a central band; even forces a choice either side of the middle | An odd scale invites central tendency. An even scale removes the hiding place but generates more boundary disagreement. |
| Symmetric or asymmetric | Equal bands both axes, or more severity bands than likelihood bands | Symmetry implies both axes deserve equal weight. Where harm to people is the concern that is usually wrong. |
| Band descriptors | Written definitions, or labels alone | Labels alone ("moderate", "possible") mean whatever the assessor wants. Undefined bands are the largest single source of inconsistency. |
| Likelihood band basis | Qualitative descriptors, or frequency ranges tied to a stated period | A descriptor without a period is meaningless. Unlikely per task, per year and per working lifetime differ by orders of magnitude. |
| Cell ratings | Product of the band numbers, or a rating assigned per cell by judgement | The product inherits the ordinal defect below. Assigning cells deliberately lets you place a severity gate. |
| Band-to-action mapping | Defined action and authority level per band, or nothing | With no mapping the rating changes no decision and the exercise is decoration. |
The finer grid is not the better grid. A 5x5 does not measure risk more accurately than a 3x3, it expresses the same judgements with more apparent resolution.
6. The honest core: the arithmetic is not real arithmetic
Here is the point that matters most and is said least. Likelihood and severity bands are ordinal. They are ranks, not quantities. Band 4 is higher than band 3, and that is the entire information content. It does not mean band 4 is one third worse than band 3, and it certainly does not mean band 4 is four times band 1.
Multiplying two ordinal labels produces a number with no meaningful units and no consistent interpretation. It looks like a calculation; it is a coincidence of notation. Three consequences follow. Wildly different risks collide on the same score while demanding completely different responses. The same one-band change in judgement moves the score by very different amounts depending on where on the scale you are, so the scheme is far more sensitive to judgement error at the top than the bottom. And equal scores are routinely not comparable, which undermines the sorted register the score was introduced to produce. A demonstration, using my own clearly invented illustrative ratings on a hypothetical five-band scheme:
| Illustrative scenario (invented) | Sev | Like | Product | What the response should be |
|---|---|---|---|---|
| A: fall from height, annual roof inspection, credible fatality | 5 | 2 | 10 | Stop and engineer out, whatever the likelihood judgement. |
| B: minor hand abrasion handling packaging, frequent | 2 | 5 | 10 | Gloves, a toolbox talk, review next cycle. Nobody stops work. |
| C: as A, one band more pessimistic on likelihood | 5 | 3 | 15 | Identical response to A. The shift moved the score by 5. |
| D: as B, one band more pessimistic on severity | 3 | 5 | 15 | Still routine. Same one-band shift, same 5 points, different meaning. |
| E: burn from infrequent hot-surface contact | 4 | 2 | 8 | Serious injury potential, yet it sorts below B by score. |
A and B share a score and are not remotely the same problem. A and C differ by a single debatable judgement and are the same problem. E has serious-injury potential and sorts below a hand abrasion. A register sorted by this column, which is how most registers reach management, puts the abrasion above the burn and buries the fatality risk mid-page. The score did not fail because somebody rated badly. It failed because the multiplication was never meaningful.
This is not a novel criticism, nor is it confined to safety matrices. It is the defect the Risk Priority Number carries in failure modes and effects analysis, where severity, occurrence and detection ratings are multiplied to rank failure modes and produce the same incomparable, discontinuous ordering. IEC 60812:2018 (Edition 3), "Failure modes and effects analysis (FMEA and FMECA)", is the international standard for that technique, and the move away from a single multiplied priority number towards structured action-priority logic is a direct response to the problem. The worked treatment is in the FMEA and RPN walkthrough.
What this does not mean
It does not mean abandon the matrix. It means stop treating the product as the answer. Use the cell, not the number: assign a rating and a required action to each cell by deliberate judgement, so the grid encodes your organisation's risk appetite rather than the accident of multiplication. And when a register goes to management, sort it by severity band first.
7. Why a severity gate matters more than any score
This is the safety-critical point on the page and I will state it without hedging. A hazard with credible fatality or serious-injury potential must not be allowed to fall below your action threshold because somebody judged its likelihood low. A competent scheme applies a severity override: above a defined severity band the rating cannot drop below a stated level whatever the likelihood judgement, and the required action and authority level follow from the severity rather than the product.
Severity is the more reliable of the two judgements, because the credible worst outcome follows from physics and from the anatomy of the person exposed. Likelihood is uncertain, contested and, as section 10 covers, systematically biased downward under pressure. A scheme in which the reliable judgement can be cancelled by the unreliable one is the design error that lets a fatality risk sit as a green line for three years. Being wrong about a low likelihood on a fatality hazard is unrecoverable; being over-cautious costs you a controls review.
The gate also connects the matrix to the rest of the management system. Where a credible serious-harm outcome is identified, what follows is not a lower score but a controls decision taken in the required order. That order is a requirement in ISO 45001:2018 (certifiable, and to be cited as amended by Amd 1:2024 on climate action) at clause 8.1.2, and in the United States in ANSI/ASSP Z10.0-2019 at section 8.4; both are voluntary documents rather than law. The principle itself is described publicly by NIOSH in the United States, which has no regulatory power of its own. A high severity band should route you into that sequence, and a low score should never route you out of it.
8. Inherent and residual risk
Most schemes rate twice. The pre-control rating is called inherent, initial or unmitigated risk; the post-control rating is residual risk. The inherent rating shows how much work the controls are doing: a hazard rated at the top of the grid before controls and comfortably low after is carried entirely by control effectiveness, which tells you where to focus verification, supervision and audit. Skip it and you lose sight of which controls are load-bearing. The residual rating answers the question the decision-maker actually asked: given what we have in place, is this acceptable.
The abuse is constant. Residual risk gets rated optimistically to bring an assessment below a threshold so it can be signed and closed, and the controls listed are the ones somebody intends to have rather than the ones verified today. The rating then documents an aspiration as though it were a state of affairs.
Two habits prevent most of that. Rate residual risk only against controls in place and verifiable now, recording planned controls separately with an owner and a date. And where the controls for a high-severity hazard are largely administrative, procedures, training, supervision and signage, be sceptical of a large drop in the rating: those controls depend on human compliance under production pressure and degrade quietly between audits. That connects to how permits and documented authorisations behave in practice, covered in permit-to-work integration.
9. This is not a criticality matrix
Two grids that look almost identical do entirely different jobs. A risk assessment matrix rates activities and hazards: a task, a job step, an operation, a situation in which somebody could be harmed, and its output drives a safety decision about whether and how work proceeds. An asset criticality matrix ranks equipment: it combines the consequence of an asset failing with something like its probability of failure to order importance across an asset register, driving maintenance strategy, spares holding, inspection frequency and capital prioritisation.
The axes differ, the population differs, and the decision differs. A high-criticality pump is not a high-risk activity, and a high-risk activity can involve an uncritical asset. If asset ranking is what you need, see the criticality matrix guide and, in its practical maintenance form, asset criticality classification. Keep the two grids separate. An organisation that runs both from one matrix ends up with a safety register full of pumps.
10. Who does the rating, and why consistency beats accuracy
Give the same hazard to two competent assessors and they will often land in different cells; give it to two teams from different sites and the spread widens. That is the normal behaviour of an instrument built from human judgement, and the implication is clear: the value of a matrix comes from a documented, shared interpretation of the bands and a recorded rationale, not from precision. In practice that means writing descriptors for every band and putting them on the form, calibrating assessors against worked examples so "unlikely" means approximately the same thing across the organisation, requiring a short written rationale because a rating with reasoning can be challenged while a bare number cannot, rating in a small group including somebody who does the job, and reviewing ratings for drift across sites.
There is a second, less comfortable effect. When a rating determines whether work can proceed, whether a job needs a permit or whether money must be spent, ratings drift downward. This is not individual dishonesty and treating it as such gets you nowhere. It is a predictable organisational response to an incentive, arriving through slightly more optimistic likelihood judgements and slightly more generous assumptions about controls. You design around it rather than appealing to integrity: severity gates remove the most dangerous target for gaming, required rationale makes optimism visible, and a cluster of ratings immediately under an action boundary is one of the more reliable audit signals I know.
The uncomfortable payoff
A matrix applied consistently but imperfectly across an organisation is far more useful than a better matrix applied differently at every site. Consistency lets you compare, trend, prioritise and audit. Accuracy, on an instrument with no units, was never available.
11. What the output must drive
A rating that does not change a decision is decoration. If a hazard rated at the top of the grid and one rated in the middle both end with the assessment filed and the job unchanged, the matrix is producing paperwork rather than protection.
A working scheme maps every band to two things: a required action, from proceed with existing controls, through implement additional controls within a defined period, to do not proceed until the risk is reduced; and an authority level, so accepting a serious risk becomes a visible act by a named person rather than a supervisor's signature. I will not publish thresholds or the actions each should trigger, because those depend on your legal duties and your organisation's risk appetite. Three notes hold generally: every band needs a defined review period, because a rating is a snapshot of conditions that change; the highest bands should carry a hard stop, since a band that can be overridden informally is not a band; and the actions should be written so a supervisor can apply them without interpretation.
12. How matrices fail in practice
The patterns are consistent enough that I look for them in the first ten minutes of reviewing any register:
- The matrix used as the assessment rather than a summary of it. A form with a hazard title, two numbers and a colour, and nothing about how the job is done, who is exposed or why the controls are adequate.
- Bands with no defined descriptors. Every assessor invents their own scale and nothing is comparable.
- Scores recorded without reasoning. Nobody can reconstruct why, challenge it, or update it. The number becomes unfalsifiable.
- Residual ratings assumed rather than verified. Controls listed on paper, credited in the rating, never confirmed in the field.
- The matrix applied to a hazard it was never designed for. An occupational-harm matrix used on a low-frequency, catastrophic-consequence process scenario, where a stronger technique from the IEC 31010 catalogue is needed.
- Copy-paste assessments. The same ratings reproduced across sites with different plant, people and conditions.
- Ratings never revisited. An assessment four years old, describing replaced equipment and a rewritten procedure, still carrying its original green rating.
None is a flaw in the concept; each is a matrix asked to carry weight it cannot bear. For how it sits inside a broader workflow see hazard identification and risk assessment and, at task-step level, the job safety analysis guide; for orientation on the wider function, what HSE actually covers.
13. A word on the legal position
One distinction gets merged constantly in training material: the duty to assess risk is a legal matter in some jurisdictions, and the matrix is not. In Great Britain the general duty to carry out a suitable and sufficient assessment sits in Regulation 3 of the Management of Health and Safety at Work Regulations 1999 (SI 1999/3242), made under the Health and Safety at Work etc. Act 1974; Northern Ireland has its own instruments with different years, so that citation is not United Kingdom-wide. It requires the assessment to be suitable and sufficient, and specifies no matrix, no band count, no scoring method and no threshold. In United States federal law there is no general regulation requiring a written task-level hazard analysis; where written assessment is required it is specific, such as personal protective equipment selection under 29 CFR 1910.132(d) in general industry, which requires a written certification identifying the workplace, the certifying person and the date, or process hazard analysis under 29 CFR 1910.119(e), which is process-level rather than task-step.
For readers in the Gulf, where I work, neither United States OSHA regulations nor Great Britain's HSE law has legal force in the United Arab Emirates; they are voluntary benchmarks carried by contracts and client specifications. The binding federal instrument is Federal Decree-Law No. 33 of 2021 on the regulation of employment relationships, administered by MOHRE, with occupational safety and health duties in Article 13, applying to the private sector including free zones but not DIFC or ADGM. In Abu Dhabi the emirate framework is ADOSH-SF, Version 4.0, administered by the Abu Dhabi Public Health Centre. Attach the jurisdiction to every duty you cite. Primary sources: ISO , IEC , HSE (Great Britain) , US OSHA , NIOSH .
The idea to walk away with
A risk assessment matrix is a device for making two judgements explicit, comparable and recorded. It is a communication and prioritisation instrument, and a good one. It is not a measurement, its multiplication is not arithmetic in any useful sense, and the number it produces cannot be trusted to order your risks correctly.
So use it for what it does well. Define the bands in writing. Say which of the three likelihood questions you are asking. Decide deliberately what happens to exposure. Assign the rating to each cell by judgement rather than by multiplication. Gate on severity so a credible fatality can never be argued down to green. Rate residual risk only against controls you have verified. Require reasoning, not just a number. Map every band to a required action and an authority level.
Final thoughts
The commonest mistake here is not choosing the wrong grid. It is believing the grid is doing more work than it is. The reasoning, the site knowledge, the conversation with the person who performs the task, and the controls decision that follows are where the safety comes from. The matrix compresses all of that into something a manager can sort and an auditor can read, and compression always loses information. Knowing what has been lost is the practitioner's skill.
If you take one action from this article, make it the severity gate. Look at your register, find every line with credible fatality or serious-injury potential, and check whether any sits below your action threshold because somebody judged the likelihood low. In most registers, some do. Those lines are the reason the instrument exists, and a score is the one thing that should never be allowed to hide them.
Disclosure
Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.
Reviewing a risk rating scheme?
Independent advisory on how risk ratings flow through permits, work orders and asset records in CMMS, CAFM and EAM systems, and how to keep a register that supports a decision rather than filing one. 22+ years across utilities, oil and gas, manufacturing, government and facility operations.
Book a conversationRelated reading: Risk assessment: a complete guide with examples, Hazard vs risk: what is the difference, Hierarchy of controls, Criticality matrix: ranking equipment by risk, FMEA example with RPN calculation, HIRA: hazard identification and risk assessment, Job safety analysis (JSA), What is HSE, Asset criticality classification.
Muhammad Abbas
CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.
Work with me