mail@mabbaz.com Abu Dhabi, UAE

Reliability Engineering · Failure Rates · Maintenance Strategy

The Bathtub Curve Explained: Equipment Failure Rates

The bathtub curve is the first picture anyone is shown in a reliability course: failure rate falling, then flat, then rising. It is a genuinely useful teaching model and it carries one implication that changes how you plan maintenance. It is also, for most industrial equipment, not an accurate description of reality. This guide explains the curve properly, then delivers the part most articles leave out.

Muhammad Abbas September 27, 2026 ~18 min read

Almost every maintenance training deck contains the bathtub curve, and almost every one presents it as a description of how equipment behaves. It is not. It is a composite teaching model assembled from three separate failure mechanisms that happen to be plottable on the same axes, and the reliability literature has been clear for decades that very few real components display all three regions within their service life. That does not make the curve useless: one of its implications is the most operationally valuable idea in the subject, and it is the one that gets skipped. This guide gives you the model, the mechanisms, the implication, then the caveat that matters more than the model.

The message up front: learn the bathtub curve for the vocabulary and for one specific insight, that a repair or overhaul returns an item to the beginning of the curve and reintroduces infant-mortality risk. Do not use it as evidence that your equipment wears out on a schedule. Time-based replacement is only defensible where a genuine wear-out region exists and you can locate it in your own failure data. Where it does not, scheduled intrusive replacement adds infant-mortality risk while removing nothing.

1. What the bathtub curve actually plots

The bathtub curve plots failure rate on the vertical axis against age or operating time on the horizontal axis, for a population of nominally identical items. The name comes from the shape: a steep drop on the left, a long flat stretch in the middle, a rise on the right. Two things about that deserve emphasis, because getting them wrong causes most of the confusion downstream.

  • It is a population curve, not an item curve. The vertical axis is a rate: how frequently failures occur among the items still surviving at that age. A single pump does not travel along a bathtub curve. A fleet of identical pumps can produce one when you aggregate their failure history by age.
  • It is a composite, not a single mechanism. The three regions come from three unrelated causes: defects introduced before service, external events during service, and cumulative physical degradation. The curve is the envelope of three overlapping processes, which is precisely why it is fragile as a general description. Change the balance between them and the shape changes completely.

The formal vocabulary for failure rate and the related dependability terms lives in IEC 60050-192:2015, Part 192 of the International Electrotechnical Vocabulary, which supersedes the older IEC 60050-191:1990 that a great deal of training material still cites. It is voluntary and paywalled, and law nowhere by itself. For the wider discipline this sits inside, see the reliability engineering guide.

2. Phase one: infant mortality, and why it is not a manufacturing problem

The left-hand region shows a decreasing failure rate. Among a fresh population, failures are relatively frequent at first and become less frequent as time passes. The mechanism is straightforward: the population contains a minority of items that were defective from the start, those fail quickly, and once they are weeded out the survivors are the sound ones. The rate falls not because items are improving but because the weak members have already gone.

The label "infant mortality" invites people to read this as a supplier quality problem, and in electronics manufacturing that is largely what it is. In plant and facilities work it very often is not. The defects that drive early failure on a real site are introduced far closer to home than the factory:

  • Genuine manufacturing defects: a weak casting, a marginal winding, a component outside tolerance. Real, but usually the smallest share on installed plant.
  • Installation and commissioning errors: incorrect wiring, wrong rotation, missing grounding, pipework strain transmitted into a pump casing, a coupling fitted without checking soft-foot.
  • Incorrect assembly: a bearing pressed on the wrong race, a seal fitted the wrong way round, a fastener torqued by feel, a shim stack out of order.
  • Wrong lubricant, or the right one in the wrong quantity: a grease incompatible with what is already in the housing, an oil of the wrong viscosity grade, overgreasing that churns and overheats, undergreasing that starves.
  • Misalignment introduced during fitting: alignment done cold and never rechecked hot, or done to a tolerance good enough for the fitter and not for the bearing.

Every item on that list except the first is inside your own control. Infant mortality in industrial equipment is mostly a workmanship and commissioning-discipline problem dressed up as a supplier problem, and that reframing is the point of the section. The fixes are unglamorous and cheap: precision alignment as a standard rather than a specialist activity, torque procedures with calibrated tools, lubricant specification enforced at the storeroom, and commissioning signed off against measured values instead of against a technician's confidence.

The test I would apply

Pull the work orders raised in the first few weeks of service on any asset your team installed or overhauled recently. If a meaningful share are defects rather than adjustments, you have an infant-mortality problem, and no change to the PM schedule will fix it. It is fixed at the point of installation.

3. Phase two: useful life, and what "random" really means

The flat middle region shows a roughly constant failure rate. Failures still happen, but their frequency does not depend on how long the item has been in service. An item that has run for a long time is no more likely to fail in the next month than one that has run for a short time. This is the region that almost every reliability calculation you will encounter quietly assumes.

"Random" is a badly chosen word here. It does not mean causeless: every one of those failures has a cause you can usually name after the event. It means the failures are not correlated with age. The population is being hit by a stream of events whose arrival does not care how old the item is, a voltage transient, a debris ingestion, an operating excursion, a control error, a mechanical shock from elsewhere, and those events arrive at a rate set by the operating environment rather than by accumulated hours.

The practical consequence is blunt. In the constant-failure-rate region, replacing an item because of its age does nothing to its failure probability. A new item drawn from the same population faces the same stream of external events at the same rate. You have spent parts and labour, and taken on the intrusion risk discussed later, to arrive at the same hazard you started with. Improving reliability here means changing the environment or the design: filtration, surge protection, operating-envelope discipline, duty control, redundancy where consequence demands it. It does not mean changing the item more often.

4. Phase three: wear-out, and the mechanisms that produce it

The right-hand rise shows an increasing failure rate. Here age genuinely does predict failure, because something physical has been accumulating. The mechanisms are the ones a materials engineer would list:

  • Fatigue: cyclic stress accumulating damage until a crack initiates and propagates. Shafts, springs, flexing pipework, structural members.
  • Erosion and abrasion: material removed progressively by a fluid stream, entrained solids or sliding contact. Impellers and casings, valve trim, chute liners.
  • Corrosion: electrochemical loss of material, thinning a wall until it can no longer carry its load. Heat-exchanger tubes, buried pipe, coastal steelwork.
  • Cumulative damage more generally: creep at temperature, insulation ageing, elastomer hardening, filter media loading, refractory spalling.

The important shared feature is that these are monotonic. Damage accumulates and does not un-accumulate in service. That is what makes age a legitimate predictor here, and it is why a wear-out region is the necessary precondition for any defensible age-based replacement task.

Wear-out mechanisms also tend to give measurable warning, which is why they are the natural home of condition monitoring. Wall thickness can be measured, insulation resistance trended, vibration signatures shift as a bearing degrades, differential pressure across a filter climbs. Where a wear-out mechanism dominates you usually have a choice between replacing on age and replacing on measured condition, and the measured-condition route is almost always the better deal. For how a mechanism is identified and catalogued in the first place, see failure modes and how to analyse them.

5. The three phases side by side

The last column is the one most versions of this table omit, and it is where the operational judgement lives.

Phase Failure rate behaviour Typical causes What actually helps What makes it worse
Infant mortality Decreasing, as defective items leave the population. Manufacturing defects, installation and commissioning errors, incorrect assembly, wrong or wrongly applied lubricant, misalignment introduced during fitting. Precision alignment standards, calibrated torque, lubricant control at the store, commissioning sign-off against measured values, run-in monitoring, factory acceptance testing. Frequent intrusive intervention, unsupervised fitting, "while we are in there" scope creep, replacing parts that were not failing.
Useful life Roughly constant, independent of accumulated age. External and operational events: transients, contamination, debris, operating excursions, control errors, mechanical shock from elsewhere. Changing the environment or the design: filtration, surge protection, operating-envelope discipline, duty balancing, redundancy where consequence justifies it. Age-based replacement, which resets nothing and spends parts, labour and intrusion risk to return to the same hazard rate.
Wear-out Increasing, because accumulated damage makes age a genuine predictor. Fatigue, erosion and abrasion, corrosion, creep, insulation and elastomer ageing, filter loading, refractory degradation. Condition monitoring against the specific mechanism, or age-based replacement where the wear-out point is located and the spread is tight; design changes to slow the mechanism. Running past the wear-out point on hope, or taking a replacement interval from a vendor figure never validated against your own duty.

No durations, rates or percentages are given deliberately: any such figure is specific to a component, a duty and an environment, and a generic number would mislead.

6. The implication that matters: repair returns the item to the start of the curve

This is the reason to learn the curve at all, and it is routinely left out of the articles that draw it. If an item's early life carries elevated failure risk because of defects introduced during assembly and installation, then any intervention that reassembles or reinstalls the item puts it back at the left-hand end of the curve. An overhaul does not simply subtract accumulated wear. It reintroduces the whole population of workmanship risks: the bearing pressed on wrong, the seal inverted, the alignment done in a hurry on a night shift, the wrong grease from an unlabelled drum, the gasket left out. The item comes back younger in accumulated wear and more fragile in introduced defects. That has three consequences for how a PM programme is written.

  • Every intrusive task has a cost that does not appear on the work order. Labour and parts are visible. The reintroduced infant-mortality risk is not, and it is charged later as an unplanned failure recorded as a random event rather than as a consequence of the maintenance that preceded it.
  • Post-intervention is the highest-risk window, not the lowest. Teams instinctively treat a freshly overhauled machine as safe and relax monitoring. The curve says the opposite: watch most closely straight after an intervention, with a deliberate run-in check rather than a return to routine.
  • Unnecessary intrusive maintenance is not neutral, it is negative. This lands hardest with managers taught that more PM is always safer. An intrusive task addressing no real failure mode does not merely waste labour; it degrades reliability by adding defect risk while removing nothing.
The question to ask of every intrusive PM task

Which specific failure mode does opening this machine address, is that mode age-related, and is the risk I remove by opening it larger than the risk I introduce by closing it again? A task that cannot answer all three is a candidate for deletion or for conversion to a non-intrusive inspection. Non-intrusive tasks do not reset the curve, which is a large part of why condition monitoring is preferable wherever it is technically possible.

There is an honest counterweight. Some interventions genuinely are necessary, and the answer is not to stop maintaining things. It is that intrusion is a cost to be justified rather than a virtue to be maximised, and that the quality of an intervention matters as much as its frequency.

7. The honest part: the bathtub curve is a teaching model, not a description of your plant

Now the caveat, which is the reason this article exists.

For a real component to trace the full bathtub shape, all three mechanisms must be present and the wear-out mechanism must become dominant within the item's actual service life. That combination is uncommon. Most components are dominated by one or two of the three, and a great many show no wear-out region at all before they are retired, replaced for obsolescence, or superseded along with the plant around them. A control card, a solenoid, a sensor, a drive, a relay: these fail, but the data generally does not show their failure rate climbing with age in any usable way.

This is not a fringe view. The reliability-centred maintenance literature and the practice analyses behind it examined large volumes of real failure data and found several distinct conditional-probability-of-failure patterns, of which the bathtub is only one, and not the most common one. That finding belongs to the classical RCM work and the aviation and industrial practice analyses that followed it rather than to any single numbered standard, and I would keep it attributed that way. The document that governs what may legitimately be called RCM is SAE JA1011_202411, "Evaluation Criteria for Reliability-Centered Maintenance (RCM) Processes", which sets the criteria an analysis must satisfy; it is not the publication of the pattern findings. Treat the multiple-pattern result as a well-established finding of literature and practice, not as a citable clause.

Where the model breaks, and what that costs you

Time-based replacement only makes sense where a clear wear-out region exists and you can locate it with enough confidence to set an interval before it. Absent a wear-out region, a scheduled replacement removes nothing, because the failure rate it is meant to pre-empt was never rising, and it adds infant-mortality risk, because the item is reassembled and reinstalled. That is the whole case against reflexive interval-based PM, and it follows directly from the curve rather than contradicting it. The failure mode you are most likely to under-manage is not wear-out; it is the defect your own team introduces at the point of intervention.

The deeper treatment of the multiple failure patterns, what each implies for task selection, and the decision logic that turns that into policy is the subject of the reliability-centred maintenance introduction, which has a dedicated section on exactly that. If this section landed, read that next. I have deliberately not reproduced its analysis here: the bathtub curve is the entry point, RCM is where the argument is worked through properly. For the strategy comparison alongside it, see preventive vs predictive vs reactive maintenance.

8. How the curve relates to failure rate and MTBF

Failure rate and mean time between failures are the quantities the curve is really about, and the relationship between them is where the model quietly constrains every calculation you do. Failure rate is the frequency of failures among surviving items, per unit of time or cycles, and the bathtub curve is by definition a plot of that quantity against age. MTBF is a mean, total operating time divided by number of failures over a period, for repairable items; MTTF is the corresponding quantity for items that are replaced rather than repaired. For the arithmetic and the traps in each, see the failure rate formula, MTBF explained and MTTF with worked examples.

Here is the connection that matters. The tidy reciprocal relationship people habitually use, MTBF as one over the failure rate, holds only when the failure rate is constant, and constant failure rate is exactly the flat middle region of the bathtub curve and nowhere else on it. In the infant-mortality region the rate is falling; in the wear-out region it is rising; in neither does a single mean summarise the behaviour usefully. Quote an MTBF across a population partly in early life and partly in wear-out and the number is arithmetically correct and practically meaningless, because it averages two opposite trends.

This is precisely why the constant-failure-rate assumption breaks on maintained plant. Maintained plant is not a population ageing quietly through a middle region. It is a population repeatedly reset to the left-hand end of the curve by overhauls, part replacements and intrusive PM, so it carries a persistent load of post-intervention early failures mixed in with the age-independent ones. The failure history then looks constant-ish for the wrong reason, and an MTBF computed from it tells you about your maintenance practice as much as about the equipment. For how these numbers combine into availability, see reliability metrics: MTBF, MTTR and availability. I publish no benchmark figures here on purpose: there is no credible universal target for failure rate or MTBF, and the only meaningful comparison is against your own baseline in your own operating context.

9. Where Weibull analysis comes in

The standard statistical tool for asking which region of the curve a failure population is actually in is Weibull analysis. At orientation depth: the Weibull distribution has a shape parameter that determines whether the modelled failure rate decreases, stays constant or increases with age. Fit the distribution to a set of times-to-failure and the fitted shape parameter tells you which of the three behaviours your data supports. It is a way of asking the data which part of the bathtub you are standing in rather than assuming.

That is as far as I will take it here, because the fitting, the treatment of suspended items, confidence bounds and the ways a Weibull fit can mislead are a separate subject with real depth. The modelling treatment, including how Weibull work feeds remaining-useful-life estimation and connects back to FMEA, is covered in from FMEA to RUL: modelling equipment failure.

The limitation nobody mentions on the slide

Weibull analysis needs a reasonable number of failures of the same mode on the same item type under comparable duty. Most facilities and plant operations do not have that. A handful of failures spread across mixed modes, duties and vintages produces a fitted shape parameter that is essentially noise, and dressing noise in a distribution does not make it evidence. If your data will not support the analysis, say so and manage on mechanism and consequence instead. That is a legitimate engineering position, not a shortfall.

10. How to tell whether your own data shows wear-out

This is a method rather than a set of figures, and it runs against most CMMS or EAM histories without new tooling.

  • Pick one item type and one failure mode. Not "pumps". One pump class, one mode, for example bearing failure on a specific duty of end-suction pump. Mixing modes is the commonest way this analysis is wasted, because a rising mode and a flat mode averaged together look flat.
  • Establish the age clock and be explicit about what it measures. Age since installation, age since last overhaul, running hours, duty cycles. These answer different questions. If the item has been overhauled, age since overhaul is usually the honest clock, precisely because of the reset effect above.
  • Collect times to failure and keep the survivors. Items that have not failed yet are data. Discarding them biases the picture toward early failure. Statisticians call these suspended or censored observations, and any competent analysis handles them explicitly.
  • Group by age band and look at the rate, not the count. Failures per surviving item per unit time in each band. Raw counts fall with age simply because fewer items remain in service, which makes almost any population look like it is improving.
  • Ask whether the trend is real or an artefact. Did the fleet change? Was a design modification made mid-life? Did failure-coding practice change? Was there an operating-regime shift? Each fakes a trend convincingly.
  • Check the mechanism agrees with the statistics. If the numbers suggest wear-out, there should be an accumulating physical mechanism: wall loss, fatigue, visible degradation on strip-down. If nobody can name it, be sceptical of the signal. Conversely, if strip-downs clearly show progressive erosion, believe the mechanism even when the dataset is too small to prove it.
  • Only then consider an interval. A time-based replacement is defensible when a mechanism is identified, the data supports a rising rate, and the spread of failure ages is tight enough that an interval set before the rise catches most of the population. If the spread is wide, condition monitoring beats any interval you can choose.

For structured, comparable failure data in the first place, ISO 14224:2016, third edition, "Petroleum, petrochemical and natural gas industries: Collection and exchange of reliability and maintenance data for equipment", is the taxonomy and data-format reference. It is written for oil and gas but its equipment taxonomy and failure-data structure are widely borrowed outside it, and adopting that structure is the single best thing you can do to make this analysis possible in two years' time. One clarification, because the confusion is common: OREDA is a proprietary, members-only reliability database, not a standard. ISO 14224 is the standard that grew out of that work. ISO 14224 is voluntary and paywalled, like the IEC and SAE documents named above.

Most organisations discover at this point that their failure coding is not good enough to answer the question. That is a finding, and acting on it is more valuable than the analysis would have been.

11. The bathtub curve is not the P-F curve

These two get conflated constantly, including in training material that should know better, and the conflation matters because it leads people to use one where the other is required. They are not variations on a theme: different axes, different subjects, different uses. The bathtub curve describes failure rate across a population over its life. The P-F curve describes the development of a single failure, from the point where it first becomes detectable to the point of functional failure. One is a population statistic about when failures occur; the other is a timeline for one failure in progress.

Dimension Bathtub curve P-F curve
Subject A population of nominally identical items. One failure developing in one item.
Vertical axis Failure rate among surviving items. Condition or resistance to failure of that item.
Horizontal axis Age or accumulated operating time of the population. Time elapsed while a specific failure develops.
Question answered Does failure probability depend on age, and if so how? How much warning do I get, and how often must I look to catch it?
Decision it supports Whether an age-based replacement or overhaul task is defensible at all. What inspection or monitoring interval a condition-based task needs.
Typical misuse Treated as proof that equipment wears out on schedule, justifying intervals that the data does not support. Treated as a population-level life curve, or an interval set equal to the P-F interval instead of comfortably shorter than it.

They are complementary rather than competing. The bathtub curve, honestly applied, tells you whether an age-based task has any basis; where it does not, the P-F curve tells you whether a condition-based task can substitute, and at what frequency. Use them in that order. The detailed treatment of the P-F interval and how to set a monitoring frequency against it is in the P-F curve explained.

The idea to walk away with

The bathtub curve is worth knowing for two things: the vocabulary of decreasing, constant and increasing failure rate, and the insight that every repair returns an item to the left-hand end of the curve. Both will change how you write a PM programme.

What the curve is not is evidence about your equipment. It is a composite teaching model, and very few real components display all three regions within their service life. Treat it as a way of naming three mechanisms, then find out which of them your own failure data actually supports. Where a wear-out mechanism is identified and locatable, an age-based task is defensible. Where it is not, scheduled intrusive replacement adds risk while removing nothing, and the better answer is condition monitoring or a design and environment change.

Final thoughts

There is a reason this model keeps being taught as though it described reality: it is intuitive, it flatters the instinct that things wear out on a schedule, and it justifies maintenance programmes that already exist. The literature has been clear for decades that most equipment does not behave this way, but that finding is harder to put on a slide than a bath-shaped line.

The advisory position I would take with any maintenance team is this. Use the curve to teach the three mechanisms and to make the case that intervention is not free. Then stop using it as an argument and start using your own data. Pick one item type and one failure mode, run the method in section ten honestly, and let the answer decide the task. The exercise usually results in fewer intrusive tasks, better-executed ones, and a calmer plant, which is the opposite of what a naive reading of the bathtub curve would have recommended.

The standards named here, IEC 60050-192:2015, SAE JA1011_202411 and ISO 14224:2016, are all voluntary and all paywalled. None is law anywhere by itself; they bind through contracts, specifications and internal policy. If you intend to cite one, buy the current edition and read its scope, which is usually narrower than the people quoting it assume.

Primary sources: IEC , ISO , SAE International .

Disclosure

Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.

Reviewing a PM programme built on age-based intervals?

Independent advisory on failure-data analysis, PM interval justification, failure coding structure and the reliability metrics that prove a change worked. 22+ years across utilities, oil and gas, manufacturing, government and facility operations.

Book a conversation

Related reading: What is reliability engineering, The P-F curve explained, Failure rate formula and calculation, MTBF explained, MTTF formula and examples, Failure modes and how to analyse them, Reliability-centred maintenance introduction, From FMEA to RUL, Reliability metrics: MTBF, MTTR and availability.

Muhammad Abbas

CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.

Work with me
MAbbaz.com
© MAbbaz.com