mail@mabbaz.com Abu Dhabi, UAE

Reliability · Condition Monitoring · RCM

P-F Curve Explained: How Failure Develops

Almost every reliability course draws the P-F curve, and almost none of them make it operationally useful. This guide explains what the curve actually claims, why the P-F interval is the single number that decides whether condition monitoring can work at all, how different detection techniques sit at different points on the same curve, and what to do when the interval is too short for any inspection you could realistically schedule.

Muhammad Abbas September 27, 2026 ~17 min read

The P-F curve is the most drawn and least used diagram in maintenance. It appears on slide four of every condition monitoring proposal, gets nodded at, and then the programme proceeds to set inspection frequencies by habit, by contract, or by whatever the previous planner typed into the system. That is a waste, because the curve is not decoration. It is a decision tool, and if you use it properly it will tell you which failure modes condition monitoring can genuinely protect you from, which detection technique to buy, how often to inspect, and, most valuably, which failure modes you should stop trying to monitor altogether.

The message up front: the P-F interval, the time between the point a failure becomes detectable and the point the item can no longer do its job, is the number that decides whether condition monitoring is a viable strategy for a given failure mode. If the interval is comfortably longer than an inspection cycle you can actually staff and fund, periodic inspection works. If it is not, no amount of sensor spend fixes it and you need a different task, redundancy, or a design change. Everything else in this article is the reasoning behind that sentence.

1. Failures develop, they rarely just happen

Start with the observation the curve rests on. Most failures are not events. They are processes. A bearing does not go from perfect to seized in an instant; the lubricant film degrades, a micro-crack forms in a raceway, the crack spalls, the spall spreads, clearances open, the shaft starts to move where it should not, heat rises, and eventually the machine can no longer pump, drive or turn to the standard its users require. The seizure is the last moment of a story that has been running for some time.

The same is true of insulation degradation in a motor winding, of corrosion under lagging, of fouling in a heat exchanger, of erosion in a valve seat, of belt fatigue, of filter blockage, and of the slow loss of efficiency in a pump impeller. In each case the item is deteriorating while still performing. It looks fine on the round sheet and it is on its way to failing.

That is the whole opportunity. If deterioration takes time, and if some measurable property of the item changes while it deteriorates, then there is a period during which you could know about the problem before it costs you anything. The P-F curve is simply a way of drawing that period and then reasoning about its length.

The corollary matters just as much, and most treatments underplay it. Some failures genuinely are near-instantaneous for practical purposes. A control card fails, a relay welds, a brittle component fractures, a lightning surge takes out a driver. There may be a physical process underneath, but nothing observable changes at a rate and in a place you could sample. For those modes the curve has no useful width, and the honest answer is that monitoring is not the answer.

2. What P and F actually mean

Two defined points, and the definitions are stricter than the casual usage suggests.

F, the functional failure, is the point at which the item can no longer perform its required function to the standard its users accept. Note what that does not say. It does not say the item stopped. A pump that still runs but can no longer deliver its rated flow against system head has functionally failed if that flow is the required function. Defining F is therefore an act of specification, not observation: you cannot locate F until someone has written down what the item is required to do and to what standard. This is exactly why functional failure definition comes early in a reliability-centred maintenance analysis, before any task is chosen.

P, the potential failure, is the point at which the developing failure first becomes detectable. It is an identifiable physical condition indicating that a functional failure is in the process of occurring. The critical word is detectable, and detectable implies a detector. P is not an intrinsic property of the deterioration; it is the intersection of the deterioration with a chosen method of looking at it. Change the method and you move P.

Condition (ability to perform the required function)

  as-new  -----------• P  deterioration first DETECTABLE
                    \
                     \  <-- the P-F INTERVAL
                      \
  failed  -----------• F  cannot meet the required standard
                       → Time

The curve is conventionally drawn as a line that falls slowly and then steepens. That shape is a useful convention rather than a measured law: deterioration does often accelerate as damage generates more damage, which is why the last part of the interval is the least comfortable place to be standing. Do not read precision into the shape. Read two things into it: there is a window, and the window closes faster at the end than at the beginning.

3. Why the P-F interval is the number that decides everything

The interval between P and F is the whole of your opportunity. Not the detection technology, not the dashboard, not the analytics. The interval.

Think about what you have to fit inside it. You have to take a measurement or make an observation that falls after P. You have to interpret it correctly and accept it, which in most organisations means someone reviews a trend and raises a work request. You have to plan the intervention: confirm the diagnosis, identify the parts, check availability, raise a purchase order if the part is not in the storeroom, allocate a competent crew, and negotiate a window with operations. Then you have to do the work. All of that has to complete before F arrives.

So the usable part of the interval is shorter than the interval itself. If organising a shutdown on a given asset takes weeks because operations will only release it during a scheduled outage, then a warning that arrives with days to spare is not a warning, it is a countdown. You have converted a surprise breakdown into an expected breakdown, which is worth something for safety and for staging parts, but it is not the benefit condition monitoring was sold on.

The test to apply

For any failure mode you are considering monitoring, ask two questions in order. Is the P-F interval long enough to fit an inspection cycle you can realistically sustain? And is what remains after detection long enough to plan and execute the repair? A yes to the first and a no to the second means condition monitoring will tell you things you cannot act on. That is a common and expensive outcome, and it is discoverable on paper before you spend anything.

4. The P-F curve is not the bathtub curve

These two diagrams get confused constantly, including in vendor material, and the confusion produces bad decisions. They answer different questions and they have different axes.

The bathtub curve is about a population. Its vertical axis is a failure rate or hazard rate, and its horizontal axis is age. It says something about how likely an item of a given age is to fail, which is why it is the diagram you reach for when you are asking whether a fixed-interval overhaul or replacement makes any sense. If failure rate does not rise with age for a given mode, an age-based task has no basis.

The P-F curve is about a single item and a single failure mode already in progress. Its vertical axis is condition, or resistance to failure, and its horizontal axis is the time since deterioration began, not the age of the item. It says nothing about whether the failure was likely. It only describes what happens once it starts.

The practical consequence is that they select different tasks. The bathtub reasoning selects scheduled restoration and discard. The P-F reasoning selects on-condition tasks, which is to say inspection and condition monitoring. An asset can easily need both kinds of reasoning applied to different failure modes on the same nameplate.

5. The detection hierarchy: techniques sit at different points on the same curve

This is the genuinely useful part of the concept and the part most treatments skip. Because P is defined by the detector, a single developing failure has as many P points as you have ways of looking at it. The deterioration is one process. The techniques intercept it at different depths.

Consider a rolling element bearing degrading in a rotating machine. Wear debris entering the lubricant, or a change in the high-frequency acoustic signature of the contact, can be present very early, while the machine is running entirely normally by every operational measure. A characteristic vibration signature emerges later, once the defect is geometrically significant enough to excite the structure at identifiable frequencies. Later still the machine begins to run warm, because friction is now doing measurable work as heat. Later still a person standing next to it can hear that something is wrong. Last of all come the symptoms that are really consequences: smoke, lubricant loss, secondary damage, an operator reporting that output has dropped.

Read that sequence again as a procurement decision, because that is what it is. Choosing a detection technique is choosing where on the curve you want to intervene, and therefore how long a P-F interval you are buying. Earlier detection means a longer interval, which means a more relaxed inspection frequency and more planning room. It also means higher cost, more skill to interpret, and more false positives to adjudicate, because early signals are weak signals.

Detection approach Roughly where it intercepts the curve Implication for inspection interval Relative cost and skill
Lubricant and wear debris analysis Very early. Wear products appear in the oil while performance is unaffected. Longest interval bought, so the most relaxed sampling frequency of any technique. Often periodic sampling is sufficient. Moderate per sample, low on-site skill, but interpretation depends on a competent laboratory and on trending rather than single results.
Airborne and structure-borne ultrasound Very early for lubrication distress, leakage and electrical discharge. Long interval, so periodic routes work well. Good for screening many points quickly. Low to moderate equipment cost, but results are operator dependent and need a consistent measurement position and technique.
Vibration analysis After ultrasound and oil for most rolling element and gear defects. Earlier than heat for most rotating faults. Shorter interval than oil or ultrasound, so a more frequent route, or permanent sensing where the mode develops quickly. High skill to diagnose rather than merely alarm. Permanent instrumentation adds capital cost but removes route labour.
Infrared thermography Mid to late for mechanical faults. Can be comparatively early for loose or degraded electrical connections, where heat is the primary symptom. Varies sharply by application. Treat electrical and mechanical thermography as separate decisions with different intervals. Moderate equipment cost. Interpretation, emissivity and load conditions at the time of survey all materially affect the result.
Process and performance parameters Mid to late. Efficiency loss, pressure differential and power draw shift once deterioration is doing real work. Interval can be effectively continuous at near-zero marginal cost if the data is already historised. Lowest marginal cost, because the instrumentation usually exists. Requires the discipline to trend it, which is rarer than the data.
Human senses on routine rounds Late. Audible noise, visible leakage, heat felt by hand, unusual smell. Short remaining interval, so rounds must be frequent to be worth anything, and the finding is usually urgent when it comes. Cheapest to deploy and the most widely available. Depends entirely on whether operators are expected to report and are listened to.
Consequential symptoms At or past F. Smoke, alarm trip, output loss, secondary damage. No usable interval. This is detection of the failure, not of the potential failure. No inspection cost, and the full cost of the failure. This is the baseline you are trying to improve on.

The ordering above is directional and mode dependent, not universal. A cracking gear tooth, an eroding valve and a degrading winding each have their own sequence. What is universal is the principle: enumerate the ways you could detect the mode, place each one on the curve relative to the others, and then choose deliberately. For the depth on what each technique can and cannot see, and how to run the programme behind it, see the condition monitoring techniques guide, which owns that ground and goes considerably further than is useful here.

6. The inspection interval rule, and why it follows

The rule is usually stated flatly: inspect at less than half the P-F interval. Stating it flatly is why people forget it. The reasoning is short and worth carrying instead.

Your inspections are a sampling grid laid over time. The P-F interval is a window of unknown position that opens somewhere on that timeline. You do not get to choose when deterioration begins, so you must assume the window can open anywhere, including immediately after an inspection you have just completed. If your inspection interval equals the P-F interval, and the window opens just after one inspection, it closes just as the next arrives. You would be relying on perfect alignment. Half the interval guarantees that at least one inspection falls inside the window no matter where it opens, because two sampling points cannot both miss a window longer than the gap between them.

Two refinements that follow from the same reasoning, and both push you further than half.

  • One inspection inside the window is the minimum, not the target. A single detection is a single data point, and a single data point is easy to dismiss as noise, a bad sensor or a bad reading. If you want a trend, which is what actually convinces an engineer and a budget holder, you need two or three readings inside the window. That means a fraction smaller than a half.
  • Detection at the last possible moment is not useful. Halving guarantees detection somewhere in the window, possibly near its end. What you need is detection with enough of the window left to plan and execute. Work backwards from the lead time your organisation actually has, not the one on the process map.

There is an asymmetry worth stating plainly. Inspecting more often than necessary wastes inspection cost, which is usually modest and always visible. Inspecting less often than necessary means the technique silently fails to protect you, and the cost of that is a breakdown that looks, on paper, exactly like a breakdown you had no way of preventing. The error is invisible, which is why it persists. If you are going to be wrong, be wrong on the frequent side, and revisit once you have evidence.

Where the rule quietly breaks

The halving rule assumes the inspection reliably detects the condition once the condition is present. Real inspections miss things: the measurement point was inaccessible that day, the machine was at a different load, the technician was new, the route was signed off without being walked. A task with a modest probability of detection per attempt needs materially more attempts than the geometry alone suggests. Before you tighten an interval, check whether your problem is frequency or whether it is inspection quality. Tightening a frequency on a task nobody performs properly buys nothing.

7. When the P-F interval is shorter than any practical inspection interval

This is the part of the concept that most material avoids, and it is the part that saves the most money.

Some failure modes develop faster than you could ever inspect. The interval is a matter of hours or minutes, or the deterioration is only detectable by a method that requires the machine to be stripped. Half of nothing is nothing. In that situation the correct conclusion is not to buy a better sensor. It is that condition monitoring is the wrong strategy for this failure mode, and continuing to pursue it is a way of feeling protected while not being protected.

What you do instead, in rough order of preference:

  • Eliminate the mode by design. Change the component, the material, the specification, the operating regime, or the upstream condition that causes it. If a filter blocks unpredictably fast because of an upstream contamination source, the durable fix is upstream, not a faster inspection round.
  • Engineer redundancy or a protective device. If the failure will always arrive without usable warning, accept that and make its arrival survivable: a standby unit with automatic changeover, a relief path, a trip that protects the expensive component by sacrificing the cheap one. Note that protective devices introduce hidden functions of their own, which then need their own failure-finding tasks, because a protective device that has quietly failed protects nothing.
  • Substitute a scheduled restoration or discard task, but only if the mode has a genuine age-related pattern to hang it on. This is where you go back to the population question and the bathtub reasoning. If there is no wear-out age, a fixed interval task is arbitrary and may make things worse by reintroducing infant mortality at every intervention.
  • Move to continuous monitoring with automated response. Continuous measurement plus a human reviewing a dashboard does not shorten your effective response time much. Continuous measurement plus an automated protective action does. Be clear which one you are buying.
  • Accept run to failure, deliberately and on the record. If the consequence is tolerable, this is a legitimate and often correct answer. The difference between a good maintenance programme and a poor one is not whether anything runs to failure. It is whether the run-to-failure decisions were made on purpose. For the consequence framing that makes this defensible, see equipment criticality analysis.

8. Long interval, short interval: worked hypothetical examples

The table below is a teaching device. Every failure mode in it is my own invention, described qualitatively, and the relative interval lengths are illustrative only. I have deliberately published no time figures, because a P-F interval attached to a piece of equipment in an article is a number someone will copy into a maintenance plan without doing their own analysis, and it will be wrong for their machine, their duty and their environment. Use the reasoning pattern, not the values.

Hypothetical failure mode (invented for teaching) Relative P-F interval What that allows Strategy the reasoning selects
A: Gradual abrasive wear of a gearbox tooth flank in a lightly loaded, clean-environment drive. Long Many inspection cycles fit inside the window, and a trend can be built before anyone has to decide anything. Periodic lubricant sampling on a relaxed cycle, with vibration as a confirmatory check once oil flags a change. Cheap, plannable, no permanent instrumentation.
B: Slow fouling of a heat exchanger surface causing progressive efficiency loss. Long The signal is already in the process data, so effectively continuous observation at no marginal instrumentation cost. Trend the performance parameters you already historise, and clean on condition rather than on a calendar. The task is an analysis task, not an inspection task.
C: Progressive loosening of a bolted electrical connection in a distribution panel under cyclic load. Medium A periodic survey can catch it, provided the survey happens under representative load. Thermographic survey at a fixed cycle, with load conditions recorded so results are comparable. Permanent temperature sensing only where the consequence justifies it.
D: Rapid propagation of a fatigue crack in a highly stressed component once initiation has occurred. Short Almost nothing. A periodic route will land inside the window only by luck. Not an on-condition task. Look at design stress, material or duty, or move to a scheduled discard if a wear-out age genuinely exists. Continuous monitoring only if it can trip automatically.
E: Sudden loss of an electronic controller output with no observable precursor. Effectively zero Nothing. There is no detectable potential failure condition to find. Redundancy, hot spare, or accepted run to failure with a stocked replacement and a documented recovery procedure. Do not instrument it.
F: Silent failure of a protective interlock that is only called upon during an abnormal event. Not applicable Nothing, because the failure has no symptoms during normal operation. This is a hidden function. A failure-finding task: periodically test the device to confirm it still works. The P-F logic does not apply and substituting a condition task here is a classic analysis error.

Notice that of six invented modes, only three select a condition-based task, and one of those three is really a data analysis task rather than an inspection. That ratio is not pessimism. It is what a disciplined analysis normally looks like, and it is the reason monitoring programmes that try to cover everything deliver thin value everywhere. The same concentration argument runs through the comparison of preventive, predictive and reactive strategies.

9. The interval belongs to the failure mode, not the asset

Here is the question I am asked most often and cannot answer: what is the P-F interval for a pump?

The question is malformed, and seeing why is the moment the concept becomes usable. A pump does not have a P-F interval. A pump has many failure modes, and each one has its own. Bearing degradation, seal wear, impeller erosion, coupling misalignment, motor winding insulation breakdown, cavitation damage, foundation or baseplate deterioration, and control or instrumentation faults are all different physical processes developing at different rates, detectable by different methods, with different consequences when they complete. Asking for one number for the pump is like asking for the average speed of everything in a city.

This has three practical consequences.

  • You cannot set a monitoring frequency for an asset. You set it for a failure mode, or for a group of modes a single technique covers. Where one route has to cover several modes with different intervals, the shortest governs, and you should know which mode is driving your frequency and what it is costing you.
  • The analysis has to be done at mode level. Which is precisely why failure mode identification comes first and task selection second. If you have not enumerated the modes, you cannot reason about intervals at all, you can only copy a frequency from somewhere. See how to identify and analyse failure modes and the structured version in the FMEA guide.
  • Vendor claims about an asset class should be read carefully. A statement that a technology gives months of warning on a machine type is a statement about one mode, usually the most favourable one, and usually under laboratory or well-instrumented conditions. Ask which mode, detected by what, and at what detection confidence.

10. The interval is a distribution, so every choice carries risk

Even for one mode on one machine, the P-F interval is not a constant. It varies with load, speed, temperature, contamination, lubrication quality, installation quality, and how far the defect had already progressed when conditions changed. The same nominal mode on two identical machines in different service will develop at different rates. Treat the interval as a distribution with a spread, not a figure.

That has a specific consequence for interval selection. If you set your inspection frequency against the average P-F interval, you are exposed on every occasion the actual interval falls below average, which is roughly half the time. The quantity to design against is the shortest interval you can reasonably expect, not the typical one. That will feel conservative and expensive, and the honest way to present it to a budget holder is as a risk position rather than a technical fact: this frequency accepts this much chance of missing the window, and here is what missing it costs.

The uncomfortable part

Most organisations do not know their P-F intervals, and many cannot find out from their own records, because the history does not distinguish failure modes, does not record when a condition was first detected as distinct from when the work was done, and closes work orders with a free-text comment instead of a coded cause. The intervals then come from engineering judgement, supplier input and published guidance, all of which are legitimate starting points and none of which is your machine. The route to better numbers is unglamorous: record the detection date and the failure mode every time, and after a few years you will have a distribution of your own. Very few sites start collecting that data, which is why very few sites ever have it.

11. Where this sits in RCM task selection

The P-F concept is not free-standing. It is one step inside a task selection logic, and using it outside that logic is how people end up monitoring things for no reason.

The sequence, compressed: define the function and the standard, define what functional failure means, identify the failure modes that cause it, understand the effects and consequences of each, and only then choose a task. When you reach task selection, an on-condition task has to pass two distinct tests. It must be technically feasible, which is where the P-F interval does its work: there is a detectable potential failure condition, the interval is long enough and consistent enough to act within, and a practical inspection frequency exists. And it must be worth doing, which is a consequence question: for safety and hidden modes it must reduce risk to a tolerable level, and for operational and economic modes it must cost less over time than the consequences it avoids.

Both tests, not one. A technically feasible task on a mode whose failure costs little is a technically feasible waste of money. That distinction between feasible and worthwhile is the discipline that keeps inspection programmes from bloating, and it is the part that gets dropped when a monitoring programme is scoped by asset list rather than by analysis. The full seven-question structure, the consequence categories and the cost of running the analysis properly are covered in the introduction to reliability-centred maintenance. For how these ideas connect forward into remaining useful life estimation and quantitative failure modelling, see from RUL to FMEA, and for the wider discipline the concept belongs to, the reliability engineering guide.

12. What the standards actually say, and what they do not

A short orientation, because people reasonably ask whether there is a document that settles any of this. There is structure, but there are no published intervals waiting for you.

On the RCM side, SAE JA1011_202411, Evaluation Criteria for Reliability-Centered Maintenance (RCM) Processes, sets the criteria a process must satisfy to be legitimately called RCM, including the required task selection logic that the P-F reasoning sits inside. It sets criteria rather than prescribing a method, so a process can conform in several different ways, and a process that fails a criterion is not RCM whatever it is marketed as. SAE JA1012_201108, A Guide to the RCM Standard, is the companion guide and is now one generation behind JA1011. On the dependability side, IEC 60300-3-11:2009 is the application guide for reliability centred maintenance. These are available through SAE International and IEC .

On the condition monitoring side there is a coherent family worth knowing, published by ISO and available through ISO . ISO 17359:2018 is the general guidelines document, the programme-level wrapper that points at the rest. ISO 13379-1:2025 covers diagnostics, meaning what is wrong and why, which is the P end of the curve. ISO 13381-1:2025 covers prognostics, meaning how long you have, which is the reasoning about the interval itself. Worth noting for anyone maintaining a document register: those 2025 editions changed the series wording from machines to machine systems, so a search on the older phrasing may not find them.

One accuracy point that gets mangled constantly in vendor material, and I will state it carefully because the detail matters. If you need vibration evaluation criteria, the ISO 10816 series has been superseded by the ISO 20816 series part by part, and that migration is not finished. ISO 10816-6:1995 for large reciprocating machines and ISO 10816-7:2009 for rotodynamic pumps remain current documents. Do not accept a blanket claim that ISO 10816 was replaced wholesale, and check the specific part you need rather than the series.

What none of these documents will give you is an interval for your machine. They are voluntary, they are paywalled, and they describe how to set up the reasoning and the programme, not what answer to arrive at. Anyone offering you a table of P-F intervals to copy is offering you someone else's machine.

The idea to walk away with

The P-F curve is a filter, not a diagram. Run every failure mode through it and it sorts your candidates into three groups. Modes with a long, consistent, detectable interval, where cheap periodic inspection at a comfortable fraction of the interval genuinely protects you. Modes with a short interval, where only continuous monitoring with an automated response, redundancy or a design change will help, and where periodic inspection is theatre. And modes with no usable interval at all, where the correct engineering answer is to stop looking and change something structural, or to accept the failure deliberately and stock the part.

Get that sorting right and a condition monitoring programme becomes small, defensible and effective. Skip it, and you get a programme scoped by asset register and budget, generating readings on modes it cannot catch and missing the ones it could have.

Final thoughts

If you take one habit from this, make it this one: whenever someone proposes monitoring something, ask which failure mode, detected by what method, over what interval, and with how much of that interval left to plan the work. Four questions, answerable on a single page, and they will resolve most monitoring decisions faster and more honestly than any trial or proof of concept.

And start recording the two dates that let you build your own intervals: the date a condition was first detected, and the date the item would have failed or did fail, against a coded failure mode. It costs nothing beyond discipline in how work is closed out. In a few years it is the only P-F data that will genuinely be about your equipment, and it is the difference between setting intervals from judgement and setting them from evidence.

Disclosure

Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.

Setting inspection intervals you can defend?

Independent advisory on failure mode analysis, condition monitoring scope, inspection frequency logic and the maintenance data structure that lets you measure P-F intervals from your own history. 22+ years across utilities, oil and gas, manufacturing, government and facility operations. No sensor vendor margins, no reseller arrangements.

Book a conversation

Related reading: What is reliability engineering, The bathtub curve explained, Failure modes and how to analyse them, FMEA guide, Equipment criticality analysis, Introduction to RCM, Condition monitoring techniques, Preventive vs predictive vs reactive, From RUL to FMEA.

Muhammad Abbas

CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.

Work with me
MAbbaz.com
© MAbbaz.com