mail@mabbaz.com Abu Dhabi, UAE

Reliability Engineering · FMEA · Weibull & RUL

From FMEA to RUL: Modelling Equipment Failure

FMEA tells you how an asset can fail. Failure modelling tells you when, and with what uncertainty. Remaining useful life is the output planners actually want. This is the bridge between the qualitative and the quantitative: how to run an FMEA properly, why the risk priority number deserves its bad reputation, what a Weibull shape parameter really tells you, why almost all maintenance history is censored, and how a remaining useful life estimate is produced and honestly bounded.

Muhammad Abbas September 25, 2026 ~14 min read

There are two separate skills hiding inside the phrase "failure analysis", and most organisations are competent at neither because they never noticed they were different. The first is qualitative: working out systematically every way a piece of equipment can stop doing what it is there to do. That is FMEA. The second is quantitative: fitting a statistical model to the failure history you already own, so that "these bearings last about eighteen months" becomes a distribution with parameters, a hazard rate, and a defensible replacement interval. Remaining useful life sits on top of the second and is only as good as the first, because a model of the wrong failure mode is worse than no model.

The message up front: FMEA decides what is worth modelling, criticality decides whether it is worth the effort, and data availability decides what can actually be modelled. Skip the FMEA and you model a mode that does not drive your downtime. Skip the criticality screen and you spend a reliability engineer's year on assets nobody cares about. Skip the honest look at your data and you produce parameters with confidence intervals so wide they contain every decision you might have made anyway.

1. Why FMEA and RUL belong in the same conversation

In most organisations these live in different departments. FMEA is a workshop artefact, produced once during commissioning or an RCM study, filed, and never opened again. Remaining useful life arrives later attached to a condition-monitoring purchase, presented as a software output rather than an engineering conclusion. The result is a predictable pair of failures. Modelling without analysis: a team fits a degradation model to vibration data and produces a remaining useful life figure without establishing which failure mode that signature corresponds to, or whether that mode is causing the unplanned outages. And analysis without modelling: a thorough FMEA identifies forty failure modes on a chiller, and every recommended task gets an interval somebody guessed.

The connection is straightforward once you see it. FMEA produces failure modes ranked by consequence. For each mode you then ask one question: do I have, or can I get, enough observed failure data on this specific mode to fit a distribution to it? Where the answer is yes, the model gives an interval or a remaining-life estimate with honest bounds. Where it is no, you fall back on judgement, manufacturer guidance and condition monitoring, and you say out loud that that is what you are doing. That distinction, between the modes you can model and the modes you can only reason about, is the most useful output of pairing the two. For the decision logic that turns a mode into a task type, the RCM introduction is the companion piece, and the failure prediction pillar covers the P-F curve in depth. This article stays on the modelling side.

2. FMEA properly: the six things every row must say

Strip away the procedural baggage and the method is simple: decompose equipment into functions, and for each function enumerate the ways it can fail and what happens when it does. The discipline is in getting six elements right.

  • Function. What the item is there to do, with a performance standard. Not "chilled water pump" but "deliver 45 litres per second at 2.8 bar differential to the AHU header". Without the standard you cannot define the failure, because failure is departure from the standard. Teams skip this step, and skipping it corrupts everything after.
  • Functional failure. The state of not meeting the standard. One function usually has several: delivers nothing, delivers less than 45 l/s, delivers at insufficient pressure. Total and partial failure have different consequences, and lumping them together loses the partial cases, which are often the expensive ones.
  • Failure mode. The specific physical event producing the functional failure. Impeller erosion. Bearing spalling. Seal face wear. Winding insulation breakdown. This is the level at which tasks are designed and distributions are fitted, which is why granularity here matters.
  • Failure effect. What happens, described concretely enough that somebody else could classify the consequence from your words alone: evidence, secondary damage, safety and environmental exposure, operational impact, repair demanded. "Pump stops" is not an effect, it is a placeholder.
  • Failure cause. Why the mode occurs. Bearing spalling from inadequate lubrication, or contamination past a failed seal, or misalignment-induced load, or electrical fluting from a variable speed drive. The same mode with different causes gets different countermeasures. Keep cause and mode separate, and do not drill to root-cause depth here: FMEA is a breadth instrument, and depth belongs in root cause analysis on the modes that turn out to matter.
  • Detection method and current controls. How the developing failure would be noticed today, and how early. Routine vibration route. Operator round. Trip on high motor current. Nothing at all. Being honest about the "nothing at all" rows is where an FMEA earns its cost.

3. The FMEA worksheet structure

There is no mandated layout. What matters is that columns flow from function to action, and that each is populated by someone qualified to populate it.

Column What goes in it Who owns it Common error
Item / functional locationAsset or sub-assembly, referenced to the CMMS hierarchy codeAsset data ownerFree text that never matches the asset register
FunctionVerb, object, performance standard, operating contextOperationsNaming the equipment instead of the duty
Functional failureEach way the standard is not met, total and partialOperations, reliabilityOnly recording total loss of function
Failure modeThe physical mechanism, at component levelTechnicianModes at system level, too coarse to act on
Failure causeWhy the mode occurs in this contextReliability engineerDrifting into root cause depth and stalling
Failure effectEvidence, secondary damage, safety, environment, operational impact, repairOperations, maintenanceOne-line effects nobody can classify later
Current detection / controlsExisting PM, condition monitoring, alarms, rounds, or nonePlannerCrediting controls not actually performed
Severity, occurrence, detectionThe three ratings, on defined scales (section 5)Operations, reliability, plannerSeverity scored from the mode; occurrence guessed; detection flattered
Action priority and ownerThe prioritisation output, plus the task, interval or design changeFacilitator, maintenance managerBlind RPN thresholds; no named owner, which is where FMEAs die
Modelling candidate flagYes / no: enough failure data to fit a distribution?Reliability engineerColumn omitted, the gap this article exists to close

That last column is not in the standard templates and I would add it to every worksheet. It costs one word per row, and it produces an uncomfortable statistic: the percentage of your high-consequence failure modes for which you have no usable quantitative evidence at all.

4. Running the workshop: who must be in the room

FMEA is a group method and it fails predictably when the group is wrong. The most common failure is a consultant producing the worksheet alone from drawings and manuals: coherent, correctly formatted, detached from how the equipment behaves on this site. The composition I would insist on is a facilitator who owns the method rather than the content, an operator who runs this equipment on shift rather than a supervisor, a technician who repairs it and therefore holds the failure mode knowledge, a reliability engineer who arrives with the CMMS extract already analysed, and a planner who owns the current-controls column honestly and will have to turn recommendations into real job plans in Maximo, SAP PM or Planon. Specialists come in for the sections that need them.

Six to eight people is the working maximum, three-hour sessions beat full days, and throughput is lower than anybody expects: anyone promising a complete FMEA on a large system in two days is scoping very shallowly. And bring the maintenance history into the room as a document, not a memory. Several years of closed work orders coded on a Problem, Cause, Action structure is the best evidence available, and it consistently contradicts recollection, which over-weights dramatic failures and forgets frequent minor ones.

The test of a finished FMEA

Hand the worksheet to a competent engineer who was not in the room. If they can read a row and independently arrive at the same consequence classification and the same recommended action, it is written well enough. If they have to ask what you meant, it will be useless in three years. That readability test matters more than the scoring.

5. Scoring, and why the risk priority number cannot be trusted

The scored form of FMEA rates each row on three ten-point scales. Severity rates the consequence of the failure effect: it is a property of the effect, not of how often it happens, and the most common error is contaminating it with frequency. If your organisation has a corporate risk matrix, the severity scale should be that matrix. Occurrence rates how frequently the cause produces the mode, and it can and should be evidenced: if you have eleven recorded bearing failures across twenty-three similar pumps over six years, occurrence is not a matter of opinion. Detection rates the likelihood existing controls reveal the developing failure in time to act, and it is the most consistently over-scored, because teams credit the monitoring that exists on paper rather than what is actually performed. A vibration route scheduled monthly and executed quarterly does not deserve a good score, nor does a technique whose interval exceeds the P-F interval of the mode it is meant to catch, a point the condition-based versus predictive comparison develops further.

Multiply the three and you get the risk priority number, a value between 1 and 1000, against which teams set a threshold and act on everything above it. It is popular because it is simple, and heavily criticised for two reasons worth understanding rather than dismissing.

It multiplies ordinal scales. The ratings are ordinal: a severity of 8 is worse than a 4 but not twice as bad, and nothing guarantees equal spacing between levels. Multiplication assumes a ratio scale, so multiplying ordinal numbers is not a defensible operation. It shows in the output: an RPN of 120 can be 10 x 4 x 3, an uncommon but potentially fatal failure, or 2 x 6 x 10, a trivial consequence that happens often and is invisible. Those are not equivalent risks and the number cannot distinguish them.

It masks high-severity items. This is the failure that actually hurts people. A mode with severity 10, occurrence 2 and detection 2 produces an RPN of 40, below any conventional threshold, and gets closed with no action: a potentially fatal mode the arithmetic has hidden, while three moderate, frequent, poorly-detected modes score above the line and consume the action budget.

What to do instead, and the cost of doing it

A severity gate first: any mode above a defined severity gets an action regardless of its other scores, assessed before any arithmetic happens. That alone fixes the worst problem and costs nothing. Then an action priority lookup table, the direction more recent SAE and automotive FMEA guidance moved: a defined high, medium or low priority for each combination of S, O and D, avoiding multiplication entirely. The honest cost is inconvenience. You lose the sortable number that made RPN easy to report upward, and you have to defend a priority rather than point at a threshold. Background on FMEA as a general method is available through ASQ .

And never treat these scores as a substitute for asset-level criticality: FMEA scoring ranks modes within a system that asset criticality classification has already selected.

6. Going quantitative: exponential and Weibull

A failure distribution states the time to failure of a population of similar items as a probability. Rather than "these bearings last about eighteen months", it says: the probability that a bearing of this type in this duty survives beyond time t is this function, with these parameters, estimated from this many observations. That buys three things a rule of thumb cannot. You can state the probability of failure before any chosen date, cost one interval against another, and attach a confidence interval that says how much your data supports the conclusion.

The exponential distribution describes purely random failure, where the probability of failing in the next hour is the same regardless of how long the item has run. One parameter, the failure rate, and one defining property: constant hazard, meaning the item does not age. It suits electronics, complex assemblies with many independent mechanisms, and failures driven by external events. It is also the distribution most maintenance reporting implicitly assumes, because mean time between failures is an exponential concept: if you report MTBF and nothing else, you have assumed constant hazard whether you meant to or not.

The consequence is sharp. If failure really is exponential, time-based replacement is useless: replacing a component that does not age gives you one with the same failure probability, at the cost of money spent and infant-mortality risk introduced. It is why the preventive maintenance strategies question of time versus meter versus condition cannot be answered without knowing the failure pattern.

The Weibull distribution is the workhorse, because it generalises the exponential and can represent increasing, constant or decreasing failure rates depending on one parameter. It has a shape parameter, conventionally beta, and a scale parameter, eta, the characteristic life at which roughly 63 percent of the population has failed. A third location parameter shifts the origin where there is a guaranteed failure-free period; most practical fits use the two-parameter form.

7. What the Weibull shape parameter tells you

Beta carries the engineering meaning, and reading it correctly is the most valuable single skill here, because beta tells you which maintenance strategy can possibly work.

Shape (beta) Hazard rate What it usually indicates Strategy implication
beta < 1DecreasingInfant mortality: installation error, commissioning defects, wrong part fitted, maintenance-induced failureTime-based replacement makes things worse. Fix installation and repair quality; check whether intrusive PM causes the failures
beta = 1Constant (the exponential case)Random failure: external events, mixed mechanisms, much electronic and control failureScheduled replacement gains nothing. Condition monitoring if degradation is detectable, otherwise run to failure with spares cover
1 < beta < 2Slowly increasingEarly wear-out: corrosion, erosion, fatigue with wide scatter, or two modes mixed in one datasetModest benefit at best from an interval. Condition-based intervention usually better. Check for mixed modes
beta near 2Linearly increasingClassic fatigue and progressive wearAge-based replacement becomes defensible and interval optimisation meaningful
3 < beta < 4Strongly increasingWell-defined wear-out: bearing fatigue, abrasive wear, filter loading, seal face wearThe strongest case for planned replacement at a calculated interval. Condition monitoring works well too
beta > 4Very steeply increasingRapid wear-out, or an artefact of too few data points or of recording replacement policy rather than failureReplace just before the knee, but verify the data: an implausibly high beta usually means you fitted your own PM schedule

An illustration, with numbers that are illustrative rather than drawn from any specific project. Suppose a fit on seal failures across a pump population returns beta of 3.4 and eta of 31 months. Beta above 3 says orderly wear-out, so a planned interval is defensible. Eta of 31 months says roughly 63 percent have failed by then, far too late to be a target; the useful figure is a lower percentile, the time by which perhaps 10 percent have failed, which for those parameters sits in the high teens of months. The choice between 10 and 20 percent is economic, driven by the ratio of planned replacement cost to failure consequence.

Now suppose the same fit returns beta of 0.8. Every conclusion reverses. There is no wear-out to get ahead of; failures are concentrated early, pointing at installation practice, seal selection, alignment during fitting, or dry-running at start-up. The correct action is not an interval, it is a review of how seals are fitted. Teams that go straight to interval optimisation without looking at beta regularly schedule replacements against an infant-mortality pattern, which increases failures because every replacement resets the asset to the most dangerous part of its life.

A mixed dataset produces a meaningless beta. If your "pump failure" records combine bearing failures, seal failures and impeller erosion, the fit averages three mechanisms and describes none. Fit one mode at a time, which is why the FMEA has to come first and why coding granularity decides whether modelling is possible at all. ISO 14224 defines the taxonomy this depends on.

8. Censored data: why your maintenance history is not what you think

This is the section most reliability introductions rush, and it determines whether your fit is honest. A failure time is right-censored when you know the item survived to a certain age but not when it failed, because it had not failed when observation ended. Every asset still running today is right-censored, as is every component replaced during a PM before it failed, every asset decommissioned while healthy, and every unit replaced as collateral during another repair. These are suspensions. Two related cases: left-censoring, where the item had already failed before you first observed it, typically a hidden failure found on inspection; and interval censoring, where you know only that failure occurred between two inspections, which is routinely recorded as if it happened on the inspection date, quietly pushing every failure time later than it really was.

Why it matters: suppose you have twenty-three pumps, and over the observation window nine bearings failed at recorded ages while fourteen are still running. The tempting analysis fits a distribution to the nine and ignores the fourteen. That is wrong in a specific direction. The survivors are disproportionately long-lived, and discarding them removes all evidence of longevity from the dataset. Your fitted characteristic life comes out too short, you conclude the population is less reliable than it is, you replace too early, and you spend money buying reliability you already had.

The correct treatment retains suspensions as partial information: each contributes the statement "this item survived at least this long". Maximum likelihood estimation handles censored data natively, and rank adjustment does the same for median-rank regression on probability plots. If a spreadsheet analysis somebody hands you has no column for suspensions, it is almost certainly biased.

The uncomfortable arithmetic of small samples

With fewer than about five or six observed failures of a single mode, a Weibull fit will still return parameters, and the confidence interval on beta will be so wide that it spans infant mortality and wear-out at once. The software does not refuse; it prints numbers. This is the most common way failure modelling goes wrong in facilities and utilities, where asset populations are small and failure records thin. If the confidence interval on beta includes 1, you have not established that the item ages, and you cannot justify a time-based interval from that data. Saying so is more professional than presenting the point estimate. The honest options are to pool across similar assets and similar duty where the engineering justifies it, to accept condition monitoring instead, or to keep collecting.

9. From hazard rate to a replacement-interval decision

The hazard rate connects a distribution to a decision: the instantaneous probability of failure in the next small interval of time, given that the item has survived to now. That conditional phrasing is what distinguishes hazard from raw failure probability, and whether it rises, falls or stays flat with age is exactly what beta encodes. Replacement resets age to zero, which is beneficial only if hazard increases with age, because only then is a new item at lower risk than the one you removed. Where hazard is flat, the swap is neutral in reliability terms and negative in cost terms; where it decreases, actively harmful.

Given an increasing hazard, the interval decision becomes an economic optimisation:

Inputs
  Cp: cost of a planned replacement (labour, part, planned downtime)
  Cf: cost of a failure (repair, secondary damage, unplanned downtime)
  Fitted distribution (beta, eta), suspensions included
  ↓
For each candidate interval T
  Probability of reaching T, and of failing before T
  Expected cost per unit time = weighted cost / expected cycle length
  ↓
Output: the T minimising expected cost per unit time, plus the curve around it

The curve matters more than the minimum. In most real cases it is flat near its optimum, meaning a wide band of intervals performs almost identically. That is good news: you do not need a precise fit to make a good decision. It also means arguing about seventeen versus nineteen months is wasted effort, while the difference between twelve and thirty-six months is real. Two boundaries: the calculation presupposes replacement is even the right task, where failure-finding, condition monitoring and redesign are alternatives the RCM decision logic evaluates properly; and the ratio of Cf to Cp drives the result strongly, while Cf is usually the number nobody has calculated. If the cost of failure is unknown, this is not a calculation, it is a sensitivity analysis, and it should be presented as one.

10. Remaining useful life: what it is and how it is produced

Everything so far has been about populations: Weibull tells you how a class of components behaves. Remaining useful life asks how much life this specific unit, in its current observed condition, has left. Precisely, it is the time from now until the item can no longer perform its function, conditional on its observed condition and its expected future operating context. Both conditions are load-bearing: remaining life under continuous full load and under intermittent light load are different numbers, so RUL without a stated duty assumption is incomplete. Three families of approach produce these estimates.

  • Physics-of-failure. An explicit engineering model of the degradation mechanism, for instance a fatigue crack-growth law, driven by measured or inferred loads and integrated forward to a defined failure criterion. It needs no failure history at all, which makes it the only option for unique or long-lived assets. Its costs are real engineering effort per mechanism, and quiet degradation when the actual failure mode differs from the modelled one.
  • Data-driven. A statistical or machine-learning model trained on condition histories from similar units, from transparent trend extrapolation to opaque recurrent networks. Its hard requirement is run-to-failure examples, which most organisations do not have, because a well-run maintenance programme deliberately prevents the very failures the model needs to learn from. That tension is real and under-discussed; the machine learning for predictive maintenance pillar goes further into it.
  • Hybrid. A physics model supplying structure, with parameters continuously updated from observed condition data, often in a Bayesian filtering framework. This is where serious industrial practice has settled, because it gets the plausibility of the first and the site-specific adaptation of the second. It is also the most demanding to build and the hardest to explain to a maintenance manager.

The right approach is determined by what you have, not by what is most sophisticated. In facilities and utilities portfolios the honest answer for many assets is that none of the three is supportable, and the correct response is condition monitoring with threshold alarms and no RUL claim.

11. Why a point estimate is dangerous

A single-number RUL is the most consequential bad habit in this field. "This motor has 47 days of remaining useful life" reads as knowledge and is closer to a coin flip with a decimal point. The uncertainty has four independent sources, none negligible: measurement, because sensors and feature extraction are imperfect; model form, because the chosen degradation form may not be right and comparing model families usually produces different answers from the same data; parameters, estimated from limited data even when the form is right; and future load, because the estimate depends on operating conditions that have not happened yet. The last alone is often larger than the other three combined, and it never appears in a vendor demonstration, because demonstrations assume steady duty. What to present instead:

  • An interval with a stated confidence level. "Remaining useful life 30 to 55 days at 80 percent confidence, assuming current duty continues." That is an engineering statement, and a planner can work with the lower bound.
  • A probability against a decision date rather than a countdown. "Probability of functional failure before the October shutdown window: approximately 15 percent." That maps onto the decision actually being made.
  • A trend in the estimate, not just its current value. An estimate that has shortened across the last four assessments is telling you something the absolute number is not.
  • A confidence qualifier tied to data adequacy. Low confidence where the model rests on few failure examples or an unverified mechanism. Engineers trust a system that admits when it does not know far more than one that is uniformly assertive, and that trust decides whether anybody acts on it.

The interval tightens as the asset approaches failure, which is the useful property: early on it is wide and mostly says "not soon". Letting people watch it narrow builds more credibility than any single accurate prediction.

12. RUL is a planning input, not an alarm

The implementation mistake I see most often is wiring remaining useful life into an alarm. A threshold is set, the estimate crosses it, a work order is generated automatically, and the team receives what looks like an instruction derived from a calculation they cannot inspect. Within months, after two or three estimates that turned out pessimistic, the alerts are being closed without action. The model may have been fine; the workflow was wrong. RUL is a scheduling input: it lets work be positioned in an existing window rather than triggering a response. The pattern that works is a few coarse bands, each attached to a planning action:

Band What it means Planning action Who acts
MonitorLower bound beyond the next two planning windowsNo action; continue monitoring at the current intervalReliability
WatchLower bound approaching the window after nextIncrease monitoring frequency, confirm the mode, check spares and lead timesReliability, stores
PlanLower bound inside the next planning windowRaise a planned work order into that window, reserve parts, book resources and accessPlanner
ExpediteLower bound before the next window; waiting carries real riskBring the work forward or intervene out of window; consider load reduction or standby transferMaintenance manager, operations
Act nowFailure probable within days, or consequence intolerableImmediate intervention; engineering sign-off if deferral is unavoidableOperations and maintenance

Two properties make this robust: it uses the lower bound rather than the point estimate, so uncertainty is handled by the design rather than hidden by it, and it is coarse enough to be right even when the model is only approximately right, which is the normal condition. The other non-negotiable is that the band transition lands in the CMMS or EAM as work in the normal queue. Whether that is Maximo, SAP PM, Hexagon EAM or Planon, the prediction has to become an ordinary planned work order that gets scheduled, executed and closed with a real outcome recorded, and the closure detail feeds back as an observation that improves both the FMEA and the model. Without that loop the apparatus is a reporting exercise.

The idea to walk away with

The chain is short and runs in one direction. FMEA identifies what could be modelled, by enumerating failure modes at a granularity fine enough to be analysed. Criticality decides what is worth modelling, because the analysis and the data collection cost real money and only consequential modes repay them. Data availability decides what can actually be modelled, and this is the constraint that binds in practice: with a handful of censored observations you can fit a distribution, but you cannot honestly conclude anything from it.

Run that chain and you get a defensible split into three groups: modes you can model quantitatively, which get calculated intervals or remaining-life estimates with stated bounds; modes with detectable degradation but insufficient failure history, which get condition monitoring and no numerical claim; and modes with neither, which get judgement, manufacturer guidance, redesign or acceptance, declared as such. That split is more valuable than any single Weibull fit, because it tells the organisation how much confidence each maintenance decision deserves. The failure to avoid is producing numbers that outrun their evidence: a risk priority number hiding a fatal mode behind arithmetic, a Weibull fit with the survivors discarded, a remaining-life countdown with no interval around it. Each is worse than the honest qualitative judgement it replaced, because it carries an authority it has not earned.

Final thoughts

If you are starting from nothing, sequence it as follows. Get your failure coding to the point where a closed work order identifies a specific failure mode rather than a general symptom, because without that no quantitative work is possible and you will not find out for two years. Run FMEA on the systems criticality has already flagged, with the right people in the room and a severity gate in place of an RPN. For each high-consequence mode, count how many clean observations you have, suspensions included. Then model the few modes where the count supports it, present the confidence intervals alongside the parameters, and be plainly honest about the rest.

That last piece of honesty is the professional skill here, and it is scarcer than the technical ability. Fitting a distribution takes minutes in any reliability tool. Knowing that the fit does not support the decision somebody wants to make with it, and saying so, is what separates reliability engineering from reliability reporting. For wider context, the predictive maintenance practitioner's guide covers how this sits inside a working operation, and the complete guide to preventive maintenance covers the programme mechanics your intervals live inside.

Need an honest read on whether your data supports the model?

Independent advisory on FMEA facilitation, failure-mode taxonomy and coding, reliability data quality, Weibull analysis and how to present remaining useful life so planners actually use it. 22+ years across CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations. No tool vendor margins.

Book a conversation

Related reading: Reliability-centred maintenance: an introduction, Predictive maintenance and failure prediction, Asset criticality classification, Failure codes: Problem, Cause, Action, Condition-based versus predictive maintenance, AI and machine learning for predictive maintenance.

Muhammad Abbas

CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.

Work with me
MAbbaz.com
© MAbbaz.com