mail@mabbaz.com Abu Dhabi, UAE

Reliability · Metrics · Maintenance Analysis

Failure Rate: Formula, Calculation and Examples

Failure rate is the simplest reliability metric to calculate and one of the easiest to misuse. The arithmetic is a division. The difficulty is everything around it: which exposure unit you divided by, whether the rate is even allowed to be treated as constant, and whether the exposure figure you used was measured or assumed. This is a practitioner's guide to the formula, the units, the worked calculations and the assumptions that decide whether the answer means anything.

Muhammad Abbas September 27, 2026 ~16 min read

Ask five people in a maintenance meeting for the failure rate of a pump class and you will get five numbers, all arithmetically correct and none comparable. One divided failures by calendar months. One divided by running hours taken from a meter nobody has read since commissioning. One counted only the failures that generated a corrective work order, and one counted every callout including the operator resets. The formula was never the problem. What differs is the exposure each person divided by, and the set of events each person was willing to call a failure. A failure rate quoted without both of those stated is not a metric, it is a number.

The message up front: the unit matters more than the number. Failures per operating hour, per calendar year, per start, per cycle and per kilometre are all legitimately called "failure rate" and none of them are interchangeable. Get the exposure definition right, state it every time you quote the figure, and be honest about whether the constant-rate assumption you are leaning on actually holds for the asset in front of you.

1. What failure rate actually is

Failure rate is the number of failures observed per unit of exposure. That is the whole concept. "Exposure" is whatever quantity of use, time or duty the item accumulated while it was at risk of failing, and the choice of that quantity is a modelling decision, not a clerical one.

The symbol conventionally used is the Greek letter lambda, which is why you will see failure rate written as lambda in datasheets, reliability models and spares calculations.

Two components have to be defined before the division is meaningful. The first is the failure count: what counts as a failure, at what boundary, and whether a repeat failure of the same item within a short window counts once or several times. The second is the exposure: the quantity of use over which those failures were counted, aggregated across however many items are in scope. Most disagreements about failure rate are really disagreements about one of those two definitions.

The formal vocabulary for dependability terms including failure rate lives in IEC 60050-192:2015, Part 192 of the International Electrotechnical Vocabulary, which supersedes the older IEC 60050-191:1990. I am describing the concepts here in practitioner language rather than quoting the standard; if you need precise normative wording for a contract or a specification, work from the published document rather than from any article, including this one. For the wider metric family, see the reliability metrics pillar.

2. The formula

The basic calculation is a single division:

failure rate = number of failures ÷ total exposure

Total exposure is summed across everything in scope. For a single item it is that item's own accumulated exposure over the observation window. For a population it is the sum of the individual exposures, which is where the term "fleet hours" or "unit hours" comes from. Ten pumps each running 1,000 hours give 10,000 pump hours of exposure, not 1,000.

The reciprocal of the result, when exposure is measured in operating time, has the units of time per failure. That is where the relationship with mean time between failures comes from, and it is the most abused identity in reliability work. Section 5 covers the conditions under which it is legitimate. For the MTBF side see the MTBF formula guide and the MTBF explainer; for non-repairable items, where the quantity is time to failure rather than time between failures, the MTTF guide.

Nothing in the formula requires exposure to be time. Time is the most common choice, and for continuously running plant usually the right one, but it is a choice.

3. Exposure units, and why the unit matters more than the number

A failure rate is a ratio, and a ratio carries its denominator with it. Strip the denominator and the number becomes uninterpretable, yet that is exactly how failure rates travel around organisations: in a slide title, in an email, in a table whose column header says only "failure rate".

The right denominator is the one that tracks the physical mechanism causing the failures. If a bearing degrades through rotation, running hours track the damage and calendar time does not. If a switchgear mechanism wears through operation, the number of operations tracks it and running hours are close to irrelevant. Pick a denominator that does not track the mechanism and the rate will move for reasons unconnected to the asset's condition.

Exposure unit When it is the right one What it misleads about
Per operating hour Continuously or near-continuously running plant where damage accumulates with running time: pumps, fans, compressors, motors under steady duty. Hides everything that happens while the asset is idle: corrosion, moisture ingress, seal drying, battery decay. An asset that fails mostly when stopped will look excellent per running hour.
Per calendar time Age and environment driven degradation, statutory and contractual reporting, and any population whose running hours are simply not measured. Conflates hard-working and idle items. Two identical units with very different duty produce the same calendar rate, so the metric cannot separate duty from reliability.
Per start or per operation Standby and intermittent equipment, switchgear, valves, actuators, and anything where the transient is harder on the item than the steady state. Says nothing about behaviour during a long run. A unit that starts reliably and then fails after four hours of load looks fine on a per-start basis.
Per cycle Fatigue and wear driven mechanisms: lifts, doors, presses, thermal cycling on plant that ramps up and down. Treats all cycles as equal. A light cycle and a full-load cycle do very different damage, so an unweighted cycle count blurs the mechanism it was chosen to track.
Per kilometre or per distance Fleet and mobile assets where distance is the best available proxy for accumulated duty. Understates urban and stop-start duty, idling hours and auxiliary equipment use. Two vehicles with identical odometers can have very different wear histories.
Per throughput unit Process plant where failures track volume handled: tonnes, cubic metres, litres pumped, items processed. Buries changes in product mix or feed quality. The rate can move because the material changed, not because the asset did.
The test I apply

Before accepting a failure rate, ask one question: if this asset ran twice as hard next year with no change in its condition, would this number move? If yes, the denominator is not tracking duty and the rate will be read as a reliability change when it is a utilisation change. That single question catches most of the calendar-time misuse I encounter.

4. Worked calculations: the same failures, different rates

The clearest way to show why the unit dominates is to hold the failure count fixed and vary the exposure definition. Every figure below is my own, invented for illustration. None of it is a benchmark, a typical value or a target, and it should not be carried into any real calculation.

Illustrative case. A population of four identical standby pumps observed over one calendar year. Over that year the population accumulated 6 recorded failures. The units each ran roughly 1,500 hours, they each started about 500 times, and the observation window was 12 months per unit.

Exposure definition Total exposure (illustrative) Failures Resulting rate
Operating hours, whole population 4 units x 1,500 h = 6,000 unit hours 6 0.001 failures per operating hour
Calendar time, whole population 4 units x 1 year = 4 unit years 6 1.5 failures per unit year
Calendar time, mistakenly not aggregated 1 year (window length used as exposure) 6 6 failures per year, a figure four times too high
Starts 4 units x 500 starts = 2,000 starts 6 0.003 failures per start
Calendar hours instead of running hours 4 units x 8,760 h = 35,040 h 6 0.00017 failures per hour, understated by roughly six times

Five rows, one set of failures, five different numbers. Three are defensible statements about the same population under different exposure definitions. Two are errors: the un-aggregated row, which inflates the rate by exactly the population size, and the calendar-hours row, which deflates it by the ratio of calendar to running hours. Both are common and neither looks wrong on a slide.

A second illustrative case makes the point for mobile assets. Take a small van fleet with 3 recorded failures over a year, an aggregate 240,000 km driven and an aggregate 9,000 engine hours. Per distance that is 1.25 failures per 100,000 km. Per engine hour it is about 0.00033 failures per hour. Both are correct, and which one you should use depends on whether the failures are driven by distance or by running time. For a fleet doing mostly urban stop-start work those two give materially different pictures. Again, my figures, illustrative only.

5. The constant failure rate assumption

This is the heart of the article. Everything convenient about failure rate, above all the reciprocal relationship with mean time between failures, rests on an assumption that is frequently unstated and frequently false.

Dividing failures by exposure gives an average rate over the observation window. Treating that average as though it applies at every moment inside the window is only legitimate if the underlying rate really was constant across it. If the rate was falling, or rising, or doing both at different times, the average is still an average but it no longer describes the item at any particular point, and projecting it forward is a modelling error rather than a rounding error.

The classic picture of how failure rate changes with age is the bathtub curve: an early region where the rate falls as manufacturing and installation defects are flushed out, a middle region where the rate is roughly flat and failures arrive more or less randomly, and a late region where the rate climbs as wear-out mechanisms take hold. The constant-rate assumption corresponds to the middle region only. That is where the reciprocal relationship with MTBF is a fair simplification, and it is the reason so much reliability mathematics quietly assumes you are in the flat section. The shape, its regions and the important caveats about how often real equipment actually follows it are covered in the bathtub curve guide, which is the companion piece to this one and worth reading alongside it.

On maintained plant the assumption often does not hold, for reasons that are specific and worth naming:

  • Maintenance itself resets and disturbs the curve. Overhauls, component renewals and rebuilds move parts of the asset back toward the early region while the rest of it carries on ageing. A maintained asset is a mixed-age population wearing one asset tag.
  • Intrusive work introduces its own failures. Infant mortality after a major intervention is real and well recognised. The rate is genuinely elevated after the work, so a window straddling the overhaul averages two different regimes.
  • Duty changes. An asset moved from base load to cycling duty, or from clean to dirty service, has a different rate afterwards. The equipment did not change; its exposure to damage did.
  • Populations are not homogeneous. Mixing units of different age, different manufacturer batch or different operating environment into one rate produces an average of dissimilar things, which looks constant mainly because the mixing smoothed it.
  • Dominant failure mode shifts. An asset whose failures were mostly electrical in its first years and mostly mechanical later has not one rate with two phases but two mechanisms with different trajectories.

So when is the reciprocal legitimate? When you have a reasonable basis for believing the rate was flat across the window: a mature population, past commissioning teething, not yet into visible wear-out, on stable duty, with no major intervention inside the window, and ideally with the failure count plotted over time so you can see it is not trending. Under those conditions, treating failure rate and MTBF as reciprocals is a fair engineering simplification and I use it without hesitation.

It is an error when any of them fail: when the window contains a rebuild, when the asset is visibly in wear-out, when the population mixes ages or duties, or when the failure count is too small to detect a trend even if one existed. The honest output then is not a single rate but a rate per period, with the periods stated.

The limitation to be honest about

In most maintenance datasets you cannot actually test whether the rate is constant. You have too few failures per asset, too short a history and too much intervention in the middle of it. That does not make the metric useless, but it does mean the constant-rate assumption is usually being adopted for convenience rather than demonstrated from evidence. Say so when you present the number, rather than letting the reciprocal arithmetic imply a rigour the data does not support.

6. Hazard rate versus failure rate

What you get from dividing failures by exposure is an average rate over a window. It is a summary statistic about a period that has already happened. It does not describe any particular instant, and it does not by itself describe an item's current risk.

What reliability engineering usually means by hazard rate is different: an instantaneous, conditional quantity, the rate of failure at a given age conditional on the item having survived to that age. The conditioning is the important part. It is a statement about the survivors at that point in life, not about everything that started. As an item moves into wear-out its hazard rate rises even though its lifetime average, computed over the whole history, may still look modest. The two numbers answer different questions and can point in different directions for the same item.

Conflating them has practical consequences. An average rate computed across an asset's whole life, then used to justify leaving a replacement interval unchanged, is exactly the error: the average is dominated by the years when the item was healthy, while the decision depends on the risk it carries now. If the question is "what interval should I set", the relevant concept is the conditional risk at the age reached. If the question is "how many spares will this population consume next year", the average over a comparable period is the right tool.

A note on rigour. Both terms have careful formal definitions in the dependability vocabulary, and I am describing concepts and the practical difference rather than reproducing or paraphrasing them. I would not assert fine normative distinctions from IEC 60050-192:2015 without the document in front of me, so if a specification or acceptance test hangs on the precise meaning, work from the published vocabulary. The point that needs no standard to support it: an average over a past window and an instantaneous conditional rate at a current age are not the same quantity and should never be swapped for each other.

7. Measuring exposure, which is where calculations actually go wrong

In advisory work the failure count is rarely the weak link. Work order history is imperfect but it exists. Exposure is where the calculation quietly breaks, because it is usually not recorded with anything like the discipline applied to failures.

  • Calendar time substituted for running hours. The most frequent error, and usually a default rather than a decision: running hours were unavailable, so the window length was used. Tolerable on continuously running plant, enormous on standby equipment, and it always understates the rate.
  • Fleet hours not aggregated. Exposure entered as the window length rather than the sum across units, inflating the rate by exactly the population size. Easy to spot once you look for it, easy to miss when you do not.
  • Meters unread, reset or wrapped. Hour meters never read, replaced without carrying the reading forward, reset at overhaul, or rolled over. Each leaves a hole in the exposure figure that is invisible in the arithmetic.
  • Standby hours counted as running hours. An asset energised but not operating accumulates a different kind of exposure. Whether it counts depends on the failure mechanism, and the choice needs to be explicit.
  • Downtime left inside the exposure. Hours when the item was out of service and could not fail should not usually count. Leaving them in dilutes the rate, which matters on assets with long outages.
  • Items entering mid-window. A unit commissioned in month nine contributes three months of exposure, not twelve. Growing populations are routinely given full-window exposure for every unit.
  • Inconsistent failure boundary. Not strictly exposure, but the same effect. If operator resets, false alarms and no-fault-found callouts drift in and out of the count between periods, the rate moves without anything physical changing.

The fix is mostly data discipline: capture meter readings as part of routine work rather than as a special exercise, define the failure boundary in writing and hold to it, and record exposure at the same granularity as the failures. Any competent maintenance system can hold readings against an asset and aggregate them across a class, so the constraint is almost never the software. Where readings genuinely do not exist, say the rate is calendar-based and stop there, rather than presenting a running-hour rate built on assumed utilisation. The meter and reading structures described in the CMMS introduction are the mechanism; consistent failure coding is what makes the numerator trustworthy.

Where formal structure helps, ISO 14224:2016, "Collection and exchange of reliability and maintenance data for equipment", is the reference worth knowing. Written for the petroleum, petrochemical and natural gas industries, it addresses exactly this problem: equipment taxonomy, failure and maintenance data definitions, and the boundary and exposure conventions that make rates from different sources comparable. Its scope is that sector rather than a universal template, but the structural thinking transfers well. It is a voluntary standard, law nowhere by itself, and paywalled.

8. A population versus a single item

Failure rate behaves differently depending on whether you are describing a fleet or an individual, and mixing the two readings is a persistent source of bad decisions.

For a population the rate is a useful aggregate. Sum the exposure across units, sum the failures, divide. It supports population-level questions well: how much corrective work this class will generate, how many spares it will consume, how it compares with another class on similar duty. The averaging is a feature here, because the question is about the aggregate.

For a single item the same arithmetic gives a much weaker statement. One asset accumulates failures slowly, so the count is small, the rate is noisy, and it is dominated by individual history: how it was commissioned, how it has been operated, what was done to it last year. Read a single-item rate as a rough indicator alongside condition information, not as a property of the item.

The transfer error runs both ways. Applying a population rate to one asset to decide whether to replace it ignores everything specific about that unit, and its condition data is almost always the better guide. Applying one item's observed rate to a whole class generalises from a sample of one. Where a population rate is being used for an individual decision, treat it as a prior to be updated by the asset's own condition and history rather than as a verdict.

9. Small samples, and why two failures tell you nothing

Failure rate is a count-based metric, and count-based metrics are noisy when counts are low. This is the dominant source of false signal in maintenance reporting.

Consider the arithmetic. A class recording 2 failures last year and 3 this year will appear in the reporting pack as a 50 percent increase in failure rate. Reverse it, 3 then 2, and the same pack shows a 33 percent improvement. Neither movement is evidence of anything. One extra event, plausibly arriving by chance or by a change in what someone chose to log, moved the headline by a third or a half.

Three consequences follow. Always publish the failure count next to the rate, because a reader who can see "2 failures" will calibrate correctly and a reader shown only a percentage change will not. Aggregate before you conclude: rolling twelve months instead of quarterly, class level instead of asset level, so the count underneath the rate means something. And resist ranking assets or sites by a rate built on a handful of events, because those league tables are mostly reordering noise and they damage trust when the bottom-ranked site turns out to have had one bad month. For a low-count population the honest statement is that the rate is indicative and the confidence low.

10. What a failure rate is legitimately used for

Having spent most of this article on constraints, it is worth being clear that failure rate earns its place. Used within its limits it supports real decisions.

Legitimate uses:

  • Spares and consumables demand. The strongest use. A population rate over a comparable period, multiplied by expected exposure, gives a defensible expected consumption for stock planning, and the averaging that weakens individual-asset statements is exactly what you want. See the spare parts and MRO inventory guide.
  • Comparing candidate equipment. Rates from comparable duty and comparable exposure definitions are a fair basis for ranking options. The comparability caveat is doing a lot of work there, but the comparison is valid when it holds.
  • Informing interval decisions. An input alongside failure mode, consequence and detectability. It informs the decision rather than making it, and if the rate is not constant the conditional risk at current age matters more than the lifetime average.
  • Feeding reliability models. Rates are inputs to availability calculations, redundancy analysis and criticality work. See the RUL and FMEA modelling guide and the reliability engineering overview.
  • Detecting large changes. A rate that doubles on a population with a decent failure count is a real signal.

Uses that do not hold up:

  • Predicting when a specific item will fail. A rate is a population statistic and does not carry a date for an individual asset.
  • Benchmarking against an external figure. Published rates carry the exposure definitions, failure boundaries, duty and environment of wherever they came from.
  • Small-sample performance management. Rates from a handful of events cannot defensibly judge a team, a contractor or a site.
  • Projecting through a regime change. A rate from before a rebuild, a duty change or a new operating mode does not carry forward past it.

A point about provenance. External failure-rate figures come variously from industry data collection efforts, manufacturer testing under stated conditions, and reliability prediction models. In the oil and gas world the best known collection is OREDA, and it is worth being precise about what that is: OREDA is a proprietary, members-only database, not a standard. ISO 14224 is the taxonomy and data-format standard that grew out of that world; the database is not a standard. Anyone citing "the OREDA standard" has the categories confused, and anyone quoting numbers from a members-only database into your business case should be asked whether the collecting conditions resemble yours.

Why there are no typical figures in this article

I have deliberately published no benchmark or typical failure rates for any equipment class, and no view on what an acceptable rate looks like. Any such figure would depend entirely on the exposure definition, failure boundary, duty, environment and maintenance regime behind it, and quoting it stripped of those would invite exactly the comparison this article argues against. The only baseline worth measuring against is your own, computed under definitions you have written down.

The idea to walk away with

Failure rate is failures divided by exposure, and the division is the easy part. Three things behind it decide whether the answer means anything: which exposure unit you chose and whether it tracks the mechanism causing the failures, whether the rate can fairly be treated as constant across the window you averaged over, and whether the exposure figure was measured rather than assumed. Get those right and failure rate is a solid working metric for spares planning, equipment comparison and interval decisions. Get them wrong and you have a number that moves for reasons unconnected to the equipment, which is worse than no number, because people will act on it.

Final thoughts

If I could change one habit in how failure rate is reported it would be this: never quote the rate without its exposure unit and its failure count. "0.001 failures per operating hour, from 6 failures over 6,000 aggregated unit hours" is a statement a reader can evaluate, argue with and act on. "Failure rate: 0.001" has lost everything that made it a metric, and that is the form in which it travels furthest.

The second habit worth changing is the reflexive reciprocal. Converting a rate to an MTBF, or back, is arithmetically trivial and conceptually loaded: it assumes a constant rate, which assumes the flat middle region of the curve, which on maintained plant with overhauls, duty changes and mixed-age populations is an assumption rather than a finding. Make it consciously and say so, and the metric will serve you well. Make it silently and it will eventually tell you something confident and wrong.

Disclosure

Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.

Reliability numbers that do not survive scrutiny?

Independent advisory on reliability metric definitions, exposure and meter data capture, failure coding discipline and the reporting structures that make the figures defensible. 22+ years across utilities, oil and gas, manufacturing, government and facility operations.

Book a conversation

Related reading: Reliability metrics: MTBF, MTTR and availability, The bathtub curve explained, MTBF formula and calculation, MTTF: formula and examples, What is reliability engineering, Spare parts and MRO inventory in a CMMS.

Primary sources: IEC for IEC 60050-192:2015 dependability vocabulary, and ISO for ISO 14224:2016 reliability and maintenance data collection. Both are paywalled voluntary standards.

Muhammad Abbas

CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.

Work with me
MAbbaz.com
© MAbbaz.com