If you have spent any time around maintenance reporting, equipment datasheets or reliability dashboards, you have seen MTBF. It appears in vendor specifications, in monthly performance packs, in tender responses and in the "reliability" tab of almost every CMMS. It looks like a simple, confident number: a quantity of hours that tells you how dependable something is. In practice it is one of the easiest metrics in the whole discipline to state correctly and misunderstand completely, and most arguments teams have about MTBF are not arguments about arithmetic at all. They are arguments about what the number is supposed to mean.
The message up front: MTBF is an average interval between failures of a repairable item, calculated from a period of observed operating time. It is not a lifetime, not a guarantee, not a warranty, and not a promise about any individual unit. Read it as a summary of a population's failure behaviour over a window of time, and it is useful. Read it as "this unit will run for that long", and you will be wrong in a way that costs money.
1. What MTBF means in plain terms
MTBF is the mean time between failures. Strip the acronym back to its three working words and it is easier to hold onto.
- Mean means average. Not typical, not minimum, not guaranteed. An average, with all the information loss that averaging always involves.
- Time between means the interval separating one failure from the next, measured in operating time rather than calendar time. An asset that sits idle half the year is not accumulating time toward its MTBF while it is idle.
- Failures means the events where the item stopped performing its required function. Which is where most disputes start, because organisations disagree about what counts as a failure.
So a plain-English reading of "this pump has an MTBF of 4,000 operating hours" is: across the population and the observation period the figure came from, the pump ran on average about 4,000 hours of duty between one functional failure and the next. That is a statement about a fleet and a time window. It is not a statement about the pump in front of you.
The formal vocabulary home for this family of terms is the international electrotechnical vocabulary on dependability, IEC 60050-192:2015, which supersedes the older IEC 60050-191:1990 that a surprising amount of published material still cites. If you want the authoritative wording rather than a practitioner's paraphrase, that is the document to buy or consult. IEC also publishes the vocabulary terms through its free Electropedia reference site, and the standards themselves are on the IEC Webstore .
The honest point about the vocabulary
A formal international vocabulary for these terms has existed for decades, and industry still cannot agree in practice on what clock starts and what clock stops. Does a failure count if the operator restarted the machine without a work order? Does standby time count as operating time? Two teams can apply the same published definition and produce different numbers because they drew the boundaries differently. Almost every argument about reliability metrics is a boundary argument wearing a mathematics costume.
2. The arithmetic definition
The calculation is genuinely simple, which is part of why the metric spread so widely. MTBF is total operating time divided by the number of failures in that time:
Three things in that expression deserve attention before you use it on anything that matters.
- Operating time, not elapsed time. The numerator is the time the item was actually running and expected to perform. Time spent shut down for a planned overhaul is not operating time. Nor, in most conventions, is the repair downtime itself, because MTBF measures the gaps between failures rather than the failures plus their repairs.
- The failure count is a definition, not an observation. The denominator only means something if you have written down what a failure is. A functional failure is the usual test: the item can no longer do what it is required to do. A degraded-but-working condition is normally a potential failure, not a failure, and counting those inflates your denominator and deflates your MTBF.
- It can be summed across a fleet. Ten identical units each running 1,000 hours give you 10,000 unit-operating-hours. If those ten units produced five failures between them, the fleet MTBF is 2,000 hours. This is how vendors generate very large MTBF numbers from short test programmes, which we will come back to.
That is the whole definition. The step-by-step mechanics, the treatment of partial periods, fleet aggregation and the common spreadsheet traps get their own treatment in the MTBF formula and how to calculate it. This article is about understanding the term rather than running the numbers.
3. A small illustrative example
The numbers below are hypothetical and chosen only to make the arithmetic visible. Treat them as illustrative, not as anything resembling a benchmark.
Imagine a single transfer pump that is monitored over a twelve-month period. It is a duty pump with a standby partner, so it does not run continuously. Over the year the run-hour meter records 5,400 operating hours. During that year the maintenance history shows three events where the pump stopped performing its function and required corrective work: a seal failure, a bearing failure and a coupling failure.
Number of functional failures = 3 (illustrative)
MTBF = 5,400 / 3 = 1,800 operating hours
So the illustrative MTBF is 1,800 operating hours. Now look at what that figure does and does not tell you. It does not tell you the three failures were evenly spaced; they might have fallen in the same fortnight after a bad repair. It does not tell you whether the seal failure was a random event or the predictable consequence of a dry run. It does not tell you how long the pump was unavailable, because downtime is a separate question answered by MTTR and by availability. And it certainly does not tell you the pump will now run 1,800 hours before its next failure.
What it does give you is a compact, comparable summary: this pump, in this duty, in this year, averaged one functional failure per 1,800 hours of running. Recalculate it next year on the same definitions and you have a trend, which is where the metric earns most of its real value.
4. Why the word "between" matters
The "between" in mean time between failures is doing real work. You can only have a time between failures if the item survives its first failure, gets repaired and returns to service. MTBF therefore belongs to repairable items: pumps, motors, compressors, chillers, vehicles, switchgear, servers, production lines.
For an item that is not repaired, there is no "between". A fuse blows once. A bearing that is discarded rather than rebuilt fails once. A sealed sensor is replaced, not fixed. For those, the correct metric is MTTF, mean time to failure, which describes the expected time from the start of service to the single failure event. That is why two apparently redundant acronyms exist, and it is not pedantry: they answer different questions and are estimated from different data structures. One is about the recurrence pattern of a repaired population, the other is about the survival time of a non-repaired one.
If you want the distinction in full, with the estimation implications, that is the subject of MTTF, its formula and examples.
Where the distinction breaks down in the field
In real asset registers the repairable and non-repairable boundary is blurry, because it depends on your spares policy rather than on physics. The same motor is repairable at one site and a throwaway at another. What I would recommend is that you record the policy alongside the asset class, so that whoever calculates the metric in three years knows which one they are entitled to calculate. Nobody does this at first, and it is exactly the missing context that makes historical reliability figures unusable later.
5. MTBF against the adjacent metrics, at a glance
MTBF rarely travels alone. Here is the quick separation between it and the metrics it is most often confused with. Each of these deserves its own treatment, and the linked articles go deeper.
| Metric | Question it answers | Applies to | Go deeper |
|---|---|---|---|
| MTBF | On average, how much operating time passes between one failure and the next? | Repairable items | MTBF formula |
| MTTF | On average, how long does an item survive until its single failure? | Non-repairable items | MTTF explained |
| MTTR | On average, how long does it take to restore the item once it has failed? | Repairable items | MTTR explained |
| Failure rate | How frequently do failures occur per unit of operating time? | Any population under observation | Failure rate |
| Availability | What proportion of required time was the item actually able to perform? | Repairable items and systems | Reliability metrics suite |
| OEE | How effectively did the equipment convert available time into good output? | Production equipment | OEE explained |
If you are trying to build a coherent metric set rather than understand one term, start with the full reliability metrics guide covering MTBF, MTTR and availability. It works through mean down time, the difference between inherent and operational availability, a worked availability calculation for series and redundant configurations, and the specific ways the whole set gets gamed in reporting. This article is the single-term explainer that sits underneath it.
6. The biggest misreading: MTBF is not a lifetime or a warranty
This is the one that matters most, and it is the reason I wanted a standalone page for the term at all. A very large number of people read an MTBF figure as a predicted service life. They are not being careless; the units invite it. If someone hands you a number measured in hours and calls it "mean time between failures", reading it as "how long the thing lasts" is an entirely natural mistake.
It is still a mistake. An MTBF of 100,000 hours does not mean a unit will run for 100,000 hours. It does not mean most units will. It does not mean any single unit will get close. What it means is that, across the population and the conditions the figure was derived from, failures occurred at an average rate of roughly one per 100,000 unit-operating-hours. One thousand units running for one hundred hours each also produces 100,000 unit-operating-hours. If exactly one of them failed, the arithmetic yields an MTBF of 100,000 hours from a test that never ran a single unit for more than about four days.
That is not a trick. It is how population reliability estimation works, and for comparing designs it is legitimate. But it demonstrates why the figure cannot be a statement about individual longevity. 100,000 hours is over eleven years of continuous running, and no honest engineer derived that number by watching something run for eleven years.
The test I use in review meetings
When someone quotes an MTBF at me, I ask one question: how many failures is that number based on, and over how much operating time? If the answer is a large operating time and a very small number of failures, the figure is a weak estimate presented with false confidence. If nobody in the room knows the answer, the figure should not be in the pack. This single question resolves more metric disputes than any amount of formula correction.
The warranty point follows directly. MTBF carries no contractual force. A warranty is a commercial commitment with a defined term, defined exclusions and a defined remedy. MTBF is a statistical summary. A supplier quoting a large MTBF has promised you nothing about your specific unit, and if you want a reliability commitment you need it written as a warranty term, an availability guarantee or a performance-based clause, not inferred from a datasheet figure. I have seen procurement teams treat a quoted MTBF as an implied durability promise during evaluation and then discover, when units failed early, that no such promise existed anywhere in the contract.
7. What MTBF assumes about failure behaviour
Behind the simple division sits an assumption that is easy to miss. Reducing a failure pattern to a single average interval implicitly treats failures as though they arrive at a roughly steady rate, independent of how old the item is. In reliability language, it assumes a constant failure rate. Under that assumption, MTBF is simply the reciprocal of the failure rate, and each is easily converted to the other.
For a great deal of industrial equipment in mid-life, that assumption is a reasonable working approximation. Random failures dominate, age is not the driver, and the average interval is informative. This is also the assumption underpinning most published reliability figures, which is worth knowing when you read one.
It fails in two directions, and both are common.
- Early-life failures. New or newly overhauled equipment often fails at an elevated rate because of installation errors, commissioning faults, manufacturing defects and workmanship. A single averaged interval spanning the first year of a new fleet mixes infant mortality in with steady-state behaviour and produces a number that describes neither.
- Wear-out failures. For components with genuine age-related degradation, failure probability rises with accumulated running time. Averaging across a wear-out period hides exactly the signal you needed, which is that the interval is shrinking. An MTBF that is stable while the underlying pattern is accelerating is a comfortable, misleading number.
This is the territory of the bathtub curve and how failure rates change over an asset's life, and the related question of how failure rate is calculated. The practical implication is straightforward: if you suspect your failure rate is not constant, MTBF alone is the wrong lens, and you should be looking at the distribution of intervals rather than their mean.
What MTBF structurally cannot tell you
An average interval contains no information about spread, sequence or cause. Two assets with identical MTBF can have completely different operational profiles, one failing predictably every few months and the other running flawlessly for a year then failing four times in a fortnight. The second is far more damaging and MTBF will not distinguish it. If a metric cannot separate those two cases, it should never be the only reliability number on a report.
8. Why small failure counts make MTBF very noisy
Because the failure count sits in the denominator, MTBF is unstable whenever that count is small. This is arithmetic rather than opinion, and it is the most underappreciated weakness of the metric in day-to-day reporting.
Take the illustrative pump from section three: 5,400 operating hours, three failures, MTBF 1,800 hours. Now suppose one more failure occurs, or suppose one of the three is reclassified on review as an operator-induced trip rather than an equipment failure. The figures move like this, all illustrative:
5,400 / 3 failures = 1,800 hours
5,400 / 4 failures = 1,350 hours
One reclassified record, and the headline reliability figure moves by half. Nothing physical changed about the pump. If you report monthly MTBF on individual assets, this is what your chart is mostly showing you: the noise of a tiny denominator, not a reliability trend. It is also why a "reliability improvement" can appear immediately after someone tightens the definition of what counts as a failure, which is the most common and least deliberate form of metric gaming.
Two practical defences. First, aggregate: calculate MTBF at asset-class or fleet level where the failure counts are large enough to mean something, and over periods long enough to accumulate them. Second, publish the denominator: report the operating hours and the failure count alongside the MTBF, always. A figure of 1,800 hours based on three failures and a figure of 1,800 hours based on ninety failures are not the same claim, and only one of them deserves to drive a decision.
9. What MTBF is genuinely good for
Having spent several sections on its limits, it is worth being clear that MTBF is a useful metric when it is used for what it is. I would not remove it from a reporting pack. I would place it carefully.
- Tracking your own trend on a stable definition. The most defensible use. Same asset class, same failure definition, same operating-time source, compared period over period. You are not claiming to know the true reliability; you are asking whether it is getting better or worse. That question is answerable and actionable.
- Comparing similar populations under similar duty. Two groups of comparable pumps in comparable service, one with a condition-monitoring programme and one without, over enough hours to accumulate real failure counts. This is one of the cleaner ways to evidence whether a maintenance strategy change worked.
- Feeding maintenance interval and spares decisions. Failure frequency is a legitimate input into how often you intervene and how many spares you hold, used alongside criticality and failure consequence rather than on its own. It also helps size the argument in a preventive versus predictive versus reactive maintenance decision, where the frequency and predictability of failure determine which strategy is rational.
- Design and specification comparison. Where two components have MTBF figures derived on a comparable basis, the comparison carries some information about relative design robustness. The caveat is the "comparable basis", which you usually cannot verify.
What it is not good for: cross-industry benchmarking, target-setting against a published figure, or any claim about how long an individual unit will run. I do not publish benchmark or target MTBF values for equipment classes, and I would treat any source that does with suspicion. Duty cycle, operating environment, installation quality, maintenance regime and the local definition of "failure" vary so much between organisations that a universal target number is not a meaningful object. The only baseline worth measuring against is your own.
10. Where the failure data comes from
An MTBF figure is only as good as the two inputs behind it, and in most organisations both are weaker than the confident decimal places suggest.
Operating time ideally comes from a run-hour meter, a control system or a historian. Where it does not, people substitute calendar time and quietly change what the metric means. An asset assumed to run continuously but actually running on a duty and standby rotation will have its operating hours overstated by a large margin, and its MTBF inflated to match.
Failure counts come from maintenance history, which in practice means the work order records in a CMMS or equivalent asset system. This is where the quality problem lives. Corrective work orders raised without failure coding, failures fixed on the shop floor and never recorded, multiple work orders raised against one event, and PM work orders miscoded as corrective all distort the count. Any maintenance system will calculate an MTBF for you; none of them can tell you whether the underlying events were classified consistently. The metric is a reporting output, not a data quality control.
For industries where this data discipline has been formalised, ISO 14224:2016, "Petroleum, petrochemical and natural gas industries: Collection and exchange of reliability and maintenance data for equipment", is the reference. It provides a standardised equipment taxonomy and a structure for recording failure and maintenance data so that reliability figures from different sources can actually be compared. Even outside oil and gas it is worth reading as a model for what disciplined failure-data capture looks like. It is available through ISO . One clarification worth making, because the two are constantly conflated: OREDA is a proprietary, members-only reliability database that grew alongside this work. It is not a standard, and citing it as one is incorrect.
11. How vendors quote MTBF, and how to read it
Vendor-quoted MTBF is a different animal from the MTBF you calculate on your own assets, and the difference is rarely spelled out on the datasheet.
A supplier figure is typically derived from one of three routes: accelerated life testing on a sample, aggregated field returns across an installed base, or a predictive calculation built up from component reliability data rather than from observed failures at all. The third is common in electronics and produces a number that has never been measured, only modelled. None of these routes is illegitimate, and all three produce a figure quoted in the same units with the same air of precision.
What I would ask before letting a quoted MTBF influence a decision:
- How was it derived? Measured, aggregated from returns, or calculated from component data. If the answer is calculated, it is a design estimate, not field performance.
- Under what conditions? Ambient temperature, duty cycle, load, cleanliness and power quality all drive failure behaviour. A figure derived in a controlled laboratory does not transfer to a rooftop plant room in a Gulf summer.
- How many failures? The same question as always. A large MTBF from a handful of failures is a weak estimate with a wide confidence interval that nobody has printed.
- What counted as a failure? Suppliers frequently count only failures returned under warranty, which is a narrower definition than "stopped doing its job in service".
- Is it in the contract? If reliability matters commercially, the commitment belongs in a warranty or availability clause. A datasheet figure is marketing material with a number on it.
Note that a large quoted MTBF and poor field reliability are not contradictory. If the quoted figure came from laboratory conditions and your installation is hot, dirty and switching frequently, both statements can be true at once. That is a conditions mismatch rather than a lie, which is precisely why the derivation question matters more than the magnitude.
12. Common MTBF misuses, and what to do instead
These are the patterns I see most often. None of them requires a statistician to fix.
| The misuse | Why it goes wrong | What to do instead |
|---|---|---|
| Reading MTBF as expected service life | It is a population failure rate expressed in time units, not an individual survival time | Use it as a failure frequency. For expected life, look at wear-out data, manufacturer replacement guidance and your own age-at-failure records |
| Treating a quoted MTBF as a durability promise | It has no contractual status and no defined remedy | Put the reliability commitment in the contract as a warranty term or availability guarantee |
| Reporting monthly MTBF per individual asset | One or two failures in the denominator make the figure swing wildly | Aggregate to asset class or fleet, over longer periods, and publish the failure count alongside |
| Using calendar hours as operating time | Inflates the numerator for anything that is not running continuously | Use run hours from the meter, control system or historian, and state the source |
| Comparing your MTBF to an external benchmark | Duty, environment and failure definitions differ too much for the comparison to mean anything | Compare against your own prior periods on an unchanged definition |
| Quoting MTBF with no MTTR or availability | Failing rarely but taking a week to fix can be worse than failing often and recovering in an hour | Report MTBF, MTTR and availability together as a set |
| Using MTBF on non-repairable items | There is no interval "between" failures for something that fails once | Use MTTF and say which one you are reporting |
| Improving MTBF by narrowing the failure definition | The number moves while reliability does not, and the trend is destroyed | Freeze the definition, document it, and restate history if you ever change it |
The last row is the one that quietly does the most damage. A metric definition changed mid-series is worse than no metric at all, because it looks like evidence of improvement. If you must change what counts as a failure, recalculate the history on the new basis and say so in the report.
The idea to walk away with
MTBF answers one question: on average, how much operating time passes between failures of a repairable item. It is total operating time divided by the number of failures, and it is a summary of a population over a window, not a prediction about the unit in front of you. An MTBF of 100,000 hours is not a promise of eleven years of service; it is a statement about an observed or modelled failure rate, and it may well have come from a short test on many units.
Used on a stable definition, aggregated to a level where the failure count is meaningful, reported alongside its denominator and paired with MTTR and availability, MTBF is one of the more useful numbers in maintenance reporting. Used as a lifetime, a warranty, a monthly per-asset chart or a benchmark against somebody else's plant, it is worse than useless, because it is confidently wrong.
Final thoughts
The reason MTBF causes so much trouble is that it is easy to calculate and hard to interpret, which is the worst combination a metric can have. Every CMMS will produce one for you on request. Very few organisations have written down what a failure is, where operating hours come from, and at what level of aggregation the figure is allowed to be quoted. Those three decisions determine whether your MTBF is a useful instrument or a decorative one, and none of them is a software question.
What I would recommend, if MTBF is about to appear on a report in your organisation: write the definition down first, on one page, and get the maintenance and operations sides to agree to it before anyone publishes a figure. Then always print the operating hours and the failure count next to the number. It looks like a small reporting convention. In practice it is the thing that stops a reliability metric from becoming a debating point, because everyone can see exactly what the figure was built from.
Disclosure
Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.
Defining a reliability metric set?
Independent advisory on failure definitions, operating-time sources, failure coding and a reliability reporting set that survives scrutiny. 22+ years across utilities, oil and gas, manufacturing, government and facility operations.
Book a conversationRelated reading: Reliability metrics: MTBF, MTTR and availability, MTBF formula: how to calculate MTBF, MTTR: meaning, formula and how to improve it, MTTF: mean time to failure, Failure rate formula and calculation, The bathtub curve explained, OEE explained, Preventive vs predictive vs reactive maintenance, What is a CMMS.
Muhammad Abbas
CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.
Work with me