mail@mabbaz.com Abu Dhabi, UAE

Reliability Metrics · Maintenance Analytics

MTTF: Mean Time to Failure, Formula and Examples

MTTF, mean time to failure, is one of the most quoted and least understood numbers in reliability work. It belongs to items you replace rather than repair, it describes a population rather than your individual unit, and the way you handle items that have not failed yet changes the answer completely. This is the arithmetic, the worked examples, and the honest limits.

Muhammad Abbas September 27, 2026 ~15 min read

A manufacturer's datasheet quotes an MTTF for a bearing. A procurement engineer reads it as a guarantee. Six weeks later the bearing fails, somebody says the figure was wrong, and a supplier relationship sours over a number that was never making the promise it was assumed to make. Some version of that conversation happens regularly. The figure was probably fine. The interpretation was not. MTTF is a simple piece of arithmetic wrapped in a set of assumptions that are rarely stated out loud, and once you can state them the metric becomes genuinely useful for spares provisioning, replacement thinking and component comparison.

The message up front: MTTF is the average operating life of a population of non-repairable items. It applies to things you throw away and replace, not things you fix. It is a statement about the population, never a prediction for the individual unit in front of you. And if you calculate it from a window of observation in which some items were still running, you have to deal with that fact honestly or your answer will be wrong in a predictable direction.

1. What MTTF actually means

Mean time to failure is the expected, or average, operating time from the moment an item is put into service until it fails, taken across a population of nominally identical items. The clock measures operating time, not calendar time, unless the item is energised continuously and the two happen to coincide. It is expressed in whatever unit the operating time is measured in: hours, running hours, cycles, kilometres, switching operations.

The defining word in that sentence is average. MTTF is a single summary statistic squeezed out of a distribution of lifetimes. Some items in the population will fail well before the mean and some well after. The mean tells you where the centre of mass of that distribution sits, and it tells you nothing at all about how wide the distribution is. Two component types can share an identical MTTF while one fails within a narrow band around the mean and the other scatters wildly from very early to very late. Those two components behave completely differently in service and you would manage them completely differently, but the MTTF figure hides the difference entirely.

The formal vocabulary for this family of terms lives in IEC 60050-192:2015, Part 192 of the International Electrotechnical Vocabulary, covering dependability. That edition supersedes IEC 60050-191:1990, so anyone still citing 60050-191 is working from a withdrawn reference. It is worth knowing the vocabulary exists, and worth being realistic that its existence has not settled practice: maintenance departments, datasheets and software reports blur MTTF, MTBF and service life freely, and a definition in a paywalled document does not stop that. The practical defence is not to win a terminology argument, it is to write down what your own organisation means by each metric before you publish a number computed from it.

2. The distinction that gives MTTF its reason to exist

MTTF exists as a separate metric from MTBF because repairable and non-repairable items are fundamentally different objects statistically, not because reliability engineers enjoy having two acronyms.

A non-repairable item has exactly one failure in its entire life. A fuse blows, a sealed bearing seizes, a lamp reaches end of life, a filter element clogs beyond service, an electronic module goes open circuit. You do not repair it; you remove it and fit a new one. Because each item fails once, there is no interval "between" failures for that item. The only thing you can average is time to failure, across many items.

A repairable item, a pump, a chiller, an air handling unit, a generator, fails, gets restored to service, and fails again. For that item you can measure the intervals between successive failures, which is what MTBF does. The full comparison, including where MTBF sits alongside MTTR and availability, is covered properly in the reliability metrics guide, and the metric itself in MTBF explained. I am deliberately not turning this page into a head-to-head, because the comparison already has a home and repeating it here would help nobody. One paragraph is all the distinction needs: if you fix it, MTBF; if you bin it, MTTF.

Framing Typical items Metric that applies Why
Non-repairable Bearings, lamps, seals, filter elements, fuses, sealed sensors, electronic modules, belts, gaskets MTTF Each item has a single failure ending its life, so there is no interval between failures to average. Only time to the one failure can be averaged, across the population.
Repairable Pumps, motors, chillers, AHUs, generators, compressors, lifts, switchgear assemblies MTBF The item is restored to service after each failure and accumulates a sequence of failures, so intervals between them exist and are the meaningful quantity.
Repairable system built from non-repairable parts A pump whose bearings and seals are replaced items Both, at different levels MTBF at the pump level for planning uptime and availability; MTTF at the bearing and seal level for provisioning spares and setting replacement intervals. Mixing the levels in one report is a common analytical error here.
Repaired-but-treated-as-new A refurbished module returned to the pool Neither cleanly Refurbishment breaks the clean model in both directions. Decide explicitly whether a refurbished unit re-enters the population with a reset clock, and record the decision, because it changes every number downstream.
The test I use

Ask what physically happens to the item when it fails. If a technician removes it and fits a replacement, it is non-repairable and MTTF is the metric. If a technician opens it, fixes something inside it and returns the same serial number to service, it is repairable and MTBF is the metric. The question is about the maintenance action, not about the price or complexity of the item.

3. The arithmetic

The calculation is deliberately unglamorous:

MTTF = total operating time accumulated across the population
          / number of items that failed

Note carefully what sits in the numerator and what sits in the denominator. The numerator is all the operating time the population accumulated, including time contributed by items that have not failed. The denominator is the count of failed items, not the count of items in the population. Those two are equal only in the special case where every item you put into service has failed by the time you do the sum. That case is rarer than people assume, and the gap between the two is the censoring problem covered further down.

The parallel arithmetic for MTBF is laid out step by step in the MTBF formula walkthrough, and the two calculations look almost identical on paper. The difference is entirely in what the population is and what the denominator counts, which is why using the right one matters more than computing either one carefully.

Three practical points about the numerator. It must be operating time, so an item sitting in a store or on a standby unit that never ran contributes nothing. It must be in one consistent unit, and mixing calendar hours for some items with running hours for others is a quiet and very common corruption. And it needs a defined observation window, because "total operating time" is meaningless without saying over what period you counted it.

4. Worked examples

Every number in the table below is a hypothetical I have made up to show the mechanics. None of them is a benchmark, a typical value, or a figure you should carry into a specification. There are no credible universal MTTF benchmarks by component class, and any figure that matters to you has to come from your own data or from a supplier prepared to say how it was derived.

Illustrative case Population and window Working Result and what it means
A. Complete data
Every item failed
10 identical sealed bearings, all run to failure within the window Failure times sum to 80,000 operating hours across 10 items. 80,000 ÷ 10 8,000 hours (illustrative). The clean case. Numerator and denominator refer to the same 10 items, so the mean is the straightforward average of ten observed lifetimes.
B. Same data, wrong denominator As case A, but the analyst divides by the number installed at the start of the year, 12, two of which were never commissioned 80,000 ÷ 12 6,667 hours (illustrative). Understated by about 17 percent because two items that contributed no operating time were counted in the denominator. Denominator discipline is not pedantry.
C. Censored, handled badly (excluded) 20 items. 6 failed inside the window; 14 were still running at the end The 6 failed items accumulated 27,000 hours. Analyst discards the 14 survivors entirely: 27,000 ÷ 6 4,500 hours (illustrative). Badly understated. Throwing away the survivors throws away only the long-lived evidence, so the sample left behind is the short-lived tail of the population.
D. Same data, survivors' time included Same 20 items. The 14 survivors accumulated 91,000 hours between them by the end of the window (27,000 + 91,000) ÷ 6 = 118,000 ÷ 6 19,667 hours (illustrative). Four times case C from the identical raw data. This is the formula applied as written, and it is a far more defensible estimate, though it carries a strong assumption discussed below.
E. Same data, survivors counted as failures Same 20 items, analyst treats end-of-window as a failure event for all 20 118,000 ÷ 20 5,900 hours (illustrative). Also wrong, in the opposite direction from case D. Items that had not failed are recorded as if they had, so every survivor's life is truncated at an arbitrary date.
F. Unit mismatch Standby pump seals: 4 items, calendar hours logged for two and running hours for two Calendar hours for a 20 percent duty item overstate its operating time roughly fivefold Inflated, by an unknown amount. Not a rounding error. This one is common in CMMS data because the system records calendar dates reliably and runtime only if somebody configured a meter.

Cases C, D and E are the point of the table. They use exactly the same observations and produce answers spanning a factor of four. The arithmetic is trivial in all three. The judgement about what to do with items that had not failed yet is where the whole result is decided.

5. MTTF is a statement about a population, not a promise about your unit

This is the point most explanations skip, and it is the one that causes the real-world arguments.

An MTTF figure describes the average behaviour of a group. It does not say that your individual bearing will last that long. It does not say that most of them will. It certainly does not say that any given unit is defective if it fails early. A component with a long MTTF can fail next week, and nothing about the MTTF figure is thereby proved wrong, because the figure never claimed otherwise. An average is compatible with an enormous amount of individual variation on both sides of it.

Two consequences follow, and they are the ones worth carrying into a supplier meeting.

  • A single early failure is not evidence against a quoted MTTF. It is one observation from a distribution. To challenge a quoted figure you need a sample of your own failures large enough to show the population behaving differently, not an anecdote, however annoying the anecdote was.
  • MTTF is not a warranty period, a service life, or a replacement interval. It is not even the point by which half the population has failed, which is the median, a different statistic that only coincides with the mean for symmetric distributions. Lifetime distributions are usually skewed, so the two genuinely differ. Quoting MTTF as "the life" of a component is the single most frequent misuse I encounter.

The useful mental reframing: MTTF answers "how many of these will I get through per thousand operating hours across my estate?" It does not answer "when should I change this one?" The first question is about a population and MTTF is well suited to it. The second is about an individual and needs condition information, a failure-mode understanding, or a life distribution with its spread intact.

What the mean throws away

Reducing a lifetime distribution to its mean discards the shape and the spread, which are usually the decision-relevant parts. Whether failures cluster tightly, arrive mostly early, or rise steadily with age changes your entire maintenance response, and MTTF is blind to all three. If a replacement decision turns on that shape, MTTF is the wrong tool and you need life-data analysis, not a single number.

6. Censoring: the problem that decides your answer

Censoring is the technical name for the situation in cases C, D and E: at the moment you stop observing, some items have not failed. You know they survived at least as long as the window, and you know nothing more about them. Their true lifetimes are unknown but bounded below. This is called right-censored data, and in MTTF work it is the normal condition rather than an edge case, because you almost never get to run an entire population to destruction before you need an answer.

The two instinctive responses are both biased, in opposite directions:

  • Excluding the survivors biases the estimate low, often dramatically. The items you discard are precisely the long-lived ones, so you keep only the early failures and compute the mean of the worst part of the population. Case C in the table shows the scale of it.
  • Treating the survivors as failures biases the estimate low as well, because every survivor's life is artificially cut short at your observation date. Case E shows this. It feels conservative and defensible, which is what makes it dangerous.

Applying the formula as written, with all operating time in the numerator and only actual failures in the denominator, as in case D, handles censoring far better than either instinct. But it is not free of assumptions. That estimator is the natural one when the failure rate is constant with age, and it quietly assumes exactly that. If the items are wearing out, so that older items are increasingly likely to fail, the estimate will tend to be optimistic. If the items suffer early-life failures, the picture changes again. The bathtub curve is the standard way of picturing those three regimes, and a component's position on it determines whether the constant-rate assumption is tolerable.

I am not going to offer you a quick correction factor, because there is not an honest one. The proper treatment is life-data analysis: fitting a lifetime distribution, commonly a Weibull, to censored data using methods built for the purpose, which give you a shape parameter describing whether the population is in early-life, random or wear-out failure, plus confidence bounds on the fitted parameters. That is a real discipline with real techniques and it is genuinely worth learning if replacement decisions carry money. The failure modelling guide goes into the censoring and Weibull depth that this page deliberately does not, and that is where to go next if this section is the reason you are here.

The minimum honest practice, if you are going to publish an MTTF from your own data: state the observation window, state how many items were in the population, state how many failed, and state how you handled the ones that did not. Four extra lines. They are the difference between a number a reliability engineer can use and a number that will be argued about.

7. Where a quoted MTTF actually comes from

For purchased components you are usually handed a figure rather than calculating one, and the figure's origin determines what it is worth. There are broadly three sources.

  • Life testing under controlled conditions. The manufacturer runs a sample until a defined number fail, at specified load, temperature and duty. This is the most direct evidence, and its relevance to you depends entirely on how close the test conditions are to your service conditions.
  • Accelerated life testing. Items are stressed beyond normal service, at raised temperature, load, voltage or cycling rate, and the observed failure times are extrapolated back to normal conditions through a model of how the stress accelerates the degradation. Legitimate and widely used, but the extrapolation model is an assumption, and if the acceleration changes the failure mechanism the extrapolation is invalid.
  • Modelled from constituent component data. Common for electronic assemblies: the assembly's figure is calculated from the failure rates of its parts and its architecture, rather than measured on the assembly. A useful design-comparison tool. It is not an observation of anything actually failing.

The questions I would put to a supplier about any quoted MTTF, in this order: was it measured or modelled? At what temperature, load and duty cycle? What was the sample size and how many items actually failed? Was it accelerated, and by what model? Does it cover the whole assembly or a sub-component? And which failure modes are in scope, because a figure derived from bearing fatigue tells you nothing about the seal that will actually put the unit out of service in your application.

A supplier who answers those questions is giving you something usable. A supplier who cannot say whether the figure was measured or modelled is quoting a marketing number, to be treated as a rough ranking device between candidate parts at best.

On the collection side, if you want your own reliability data comparable across sites or with anyone else's, ISO 14224:2016 is the reference for collecting and exchanging reliability and maintenance data for equipment. Its taxonomy was written from an oil, gas and petrochemical angle, but the discipline it imposes on what you record against a failure applies well beyond that sector. One clarification worth making because it circulates as an error: OREDA is a proprietary, members-only reliability database, not a standard. ISO 14224 is the standard that grew out of that work.

8. MTTF and failure rate: the reciprocal, and its condition

You will constantly see MTTF and failure rate presented as reciprocals:

failure rate = 1 / MTTF
MTTF = 1 / failure rate

That relationship is exact, and it holds only under the assumption that the failure rate is constant with age. State the assumption whenever you use the relationship, because it is doing real work and it is frequently untrue.

A constant failure rate means an item is no more likely to fail in its next hour of operation after ten thousand hours than it was in its first: failures arrive randomly, driven by external events rather than accumulated wear. That is a reasonable model for some electronic components. It is a poor model for anything that visibly wears out, and most mechanical items wear out: seals harden, belts fatigue, filters load up, bearings accumulate damage. For those, the failure rate rises with age, so a single MTTF-derived failure rate understates the risk on old units and overstates it on new ones.

The consequence is that converting a wear-out component's MTTF into a constant failure rate, then feeding that rate into an availability or system reliability calculation, produces a confident-looking result built on a false premise, usually optimistic about exactly the old units that are about to hurt you. The mechanics of the conversion and where it breaks down are in the failure rate formula guide.

The one-line version

MTTF and failure rate are reciprocals when the failure rate does not change with age. For a wearing part that condition does not hold, so treat the conversion as a rough screening estimate and say so in the footnote rather than building a business case on it.

9. What MTTF is genuinely good for

Having spent several sections on its limits, it is worth being clear that MTTF earns its place. Three uses where I would reach for it without hesitation.

  • Spares provisioning. This is the strongest use, because it is inherently a population question, which is exactly what MTTF answers. If you know roughly how many operating hours your estate accumulates on a given part type per year, and you have a defensible MTTF for that part, you have a first-pass expected annual consumption. Add lead time and criticality on top and you have a stocking policy. The wider mechanics of translating that into holdings and reorder points sit in the spare parts and MRO inventory guide.
  • Comparing candidate components. Two seals quoted under comparable test conditions with comparable derivation are legitimately rankable on MTTF. The comparison is only as good as the comparability, which is why the supplier questions above matter, but as a screening filter at specification stage it is sound.
  • Sanity-checking replacement-interval thinking. If a fixed replacement interval sits far above the component's MTTF, you are running a large fraction of the population to failure whether you intended to or not. If it sits at a small fraction of the MTTF, you are scrapping a lot of serviceable life. MTTF will not set the interval for you, because that needs the distribution shape and the consequence of failure, but it will tell you quickly if the current interval is nowhere near sensible.

Notice what those three have in common. Each one is a question about a group of items over a period, answered to an order of magnitude, used to inform a decision rather than to make it. That is the register in which MTTF works well.

10. What MTTF cannot tell you

The corresponding list of things the metric genuinely cannot do, which is where most of the trouble starts:

  • When your specific unit will fail. No population statistic can do this. For an individual unit you need condition information, which is the domain of condition monitoring and remaining useful life estimation, not MTTF.
  • How much variation to expect. The mean carries no information about spread. Two parts with identical MTTF can behave entirely differently in service.
  • Whether the item is wearing out. That is a question about the shape of the failure rate over age, and a single mean cannot answer it in either direction.
  • Anything about consequence. A short-lived cheap filter and a long-lived critical seal are not comparable on MTTF alone, because the metric is silent on what happens when the failure occurs. Criticality and failure-mode analysis carry that information.
  • How the part behaves in your conditions. A quoted MTTF is conditional on the test environment. Duty cycle, ambient temperature, contamination, vibration, load and installation quality all shift real service life, sometimes by a lot.
  • Whether your own calculated figure is trustworthy. An MTTF computed from a CMMS whose failure records are patchy, whose runtime meters are absent, and whose replacement transactions are logged inconsistently is arithmetic performed on noise. Data quality bounds the metric absolutely, and no statistical method recovers information that was never recorded.

That last one is the constraint I run into most. MTTF from your own history requires you to know which items were fitted when, how long each ran, and whether a removal was a failure or a precautionary change. Many CMMS installations cannot answer the third question at all: a replacement is logged as a replacement, with no indication of whether the old part had failed. Where that is the case, fixing the recording is the higher-value project and it has to come before the metric. The broader discipline this sits inside is sketched in the reliability engineering guide.

The idea to walk away with

MTTF is total operating time across a population divided by the number of items that failed, applied to items you replace rather than repair. That is the whole of the arithmetic. Everything that makes it hard is interpretation: it is a population statement and not a prediction for your unit, it discards the spread that usually matters most, it is decided in practice by how you handle items that had not failed when you stopped looking, and its reciprocal relationship with failure rate holds only if the failure rate is constant with age.

Used as an order-of-magnitude population input to spares provisioning, component screening and replacement-interval sanity checks, it is a good metric that repays the small effort of computing it properly. Used as a promise about an individual item, a warranty period or a service life, it will mislead you, and it will do so with the false authority that any single number carries.

Final thoughts

If you take one operational habit from this page, make it the four-line footnote. Whenever you publish an MTTF, state the observation window, the population size, the number of failures, and how survivors were handled. That footnote makes the figure auditable, makes disagreements about it productive, and forces you to confront the censoring question before somebody else does.

And when a datasheet MTTF lands on your desk, resist the urge to either believe it or dismiss it. Ask how it was derived. Measured or modelled, at what conditions, on what sample, covering which failure modes. The answers tell you whether you are holding evidence or a marketing figure, and that distinction is worth more than any amount of care applied to the arithmetic afterwards. For the formal terminology, IEC publishes IEC 60050-192:2015, and ISO publishes ISO 14224:2016 for reliability data collection. Both are paywalled, and both are worth having on the shelf if reliability data is part of your job rather than an occasional visitor to it.

Disclosure

Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.

Building reliability metrics you can defend?

Independent advisory on metric definitions, failure and replacement data capture in CMMS and EAM, and the spares provisioning logic that sits on top. 22+ years across utilities, oil and gas, manufacturing, government and facility operations.

Book a conversation

Related reading: Reliability metrics: MTBF, MTTR and availability, MTBF explained, How to calculate MTBF, Failure rate formula and calculation, The bathtub curve explained, From RUL to FMEA: modelling equipment failure, Spare parts and MRO inventory in a CMMS, What is reliability engineering.

Muhammad Abbas

CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.

Work with me
MAbbaz.com
© MAbbaz.com