mail@mabbaz.com Abu Dhabi, UAE

Reactive Maintenance · Strategy · Reliability

Reactive Maintenance: A Complete Guide

Reactive maintenance is usually described as a failure of discipline. That is half true and the half that is missing is the important half. This guide treats reactive maintenance as what it really is: an operating posture, not a work type. It covers the terminology mess, the narrow set of conditions under which running an item to failure is genuinely the correct engineering answer, how to recognise an organisation trapped in a reactive posture, and the sequence that actually gets one out.

Muhammad Abbas September 27, 2026 ~21 min read

Walk onto a site that is running reactively and you can feel it before anybody tells you. The radios do not stop. The planner is standing in the workshop rather than sitting with next week's schedule. Everybody is working extremely hard and the reliability of the estate is not improving. That pattern recurs across CMMS, EAM and CAFM engagements, and reactive maintenance is rarely a decision anybody made. It is a state an organisation drifts into, and it is self-reinforcing once it arrives. That is what makes it worth a guide of its own, separate from the mechanics of the repair work itself.

The message up front: reactive maintenance is a posture, not a work category. Running an item to failure on purpose, after you have analysed it and concluded that no proactive task is worth doing, is a legitimate and often correct strategy. Being reactive because nobody ever made that analysis is a management deficit wearing the same overalls. From the shop floor the two look identical. In every respect that matters they are opposites, and telling them apart is the whole point of this article.

1. The terminology problem, dealt with first

Four terms get used interchangeably in the trade and they are not the same thing: reactive maintenance, run to failure, breakdown maintenance and corrective maintenance. The confusion is not pedantic. It causes organisations to report a reactive percentage that means nothing, to defend a genuine run-to-failure decision as if it were a lapse, and to excuse a genuine lapse as if it were a decision. Separate them once, properly, and the rest of the subject becomes much easier to reason about.

Reactive maintenance is a posture. It describes the state of an organisation, or of the management regime applied to a particular part of the estate, in which action follows the event because no decision was ever taken to act ahead of it. The defining feature is absence: absence of analysis, absence of a task, absence of a plan. You did not choose to wait for the failure; you simply had not decided anything else.

Run to failure is a decision. Someone identified the item, identified its credible failure modes, considered the consequences of each, evaluated the candidate proactive tasks, and concluded that none of them was worth doing, so the item will be operated until it fails and then repaired or replaced. That conclusion is documented, it has an owner, and it is revisited when the duty or the consequence changes. It is an output of analysis, not the absence of one.

Corrective maintenance and breakdown response are work categories. They describe the work that gets executed, not the regime that produced it. Corrective maintenance is the restoration of an item to a condition in which it can perform its required function, and European maintenance terminology in EN 13306:2017 (a CEN standard, with no ISO twin) further splits corrective work into immediate and deferred variants, which is exactly why "corrective" cannot be treated as a synonym for "unplanned". Breakdown response is the narrower case of reacting to an item that has already stopped functioning. Both of those are downstream of a strategy decision, whichever way that decision went, and both are covered properly in their own guides: see corrective maintenance as a work category and breakdown maintenance and the response to a breakdown event. I will not re-teach either here.

There is a fourth axis that people fold into this conversation and should not: whether the work was planned or unplanned. That is a separate dimension from whether it was proactive or reactive, and treating them as one variable is a common analytical error in maintenance reporting. The two-axis untangling lives in planned versus unplanned maintenance.

Term What it describes Is it a decision? Where it is owned
Reactive maintenance An operating posture. Action follows the event because nothing else was decided. No. It is the absence of a decision. Maintenance leadership. It is a management state, not a work record.
Run to failure A deliberate strategy for a specific item and failure mode: operate to failure, then restore. Yes. It is the documented output of analysis. Reliability or asset engineering, with a named owner and a review date.
Corrective maintenance A work category: restoring function after a fault is found. Immediate or deferred. Neither. It is the work that results from a decision. Planning and execution. Recorded as a work order type.
Breakdown response The narrower event case: an item has stopped, and the response begins. Neither. It is an event and the reaction to it. Operations and the on-call or emergency process.
The distinction that matters most

Run to failure and reactive maintenance are indistinguishable from the shop floor and opposite in quality of management. In both cases a technician is called to a stopped item and repairs it. In one case that outcome was predicted, priced, resourced and accepted. In the other it was a surprise. One is a decision, the other is the absence of one, and no amount of work-order coding will tell you which you have. Only the analysis record will.

2. Where reactive sits among the strategies, briefly

A short placement, because the full three-way comparison belongs elsewhere and I would rather you read it there than get a thinner version here. Preventive work is driven by a predetermined interval or by condition, predictive work is a form of condition-based work that estimates how long you have, and reactive work happens after the fact. EN 13306:2017 gives the formal European vocabulary for that family, and IEC 60050-192:2015, the dependability part of the International Electrotechnical Vocabulary, is the formal home of the availability and repair-time terminology that the arguments about reactive work tend to hinge on. For the strategy comparison itself, with the per-asset decision framework, read preventive versus predictive versus reactive maintenance, and for the proactive side in depth, the complete guide to preventive maintenance.

The one point I want to add rather than repeat is this: a strategy ladder implies reactive sits at the bottom and that progress means climbing off it. That framing is useful for organisational maturity and misleading for individual items. The correct strategy for a given item and failure mode is whichever one delivers the required performance at the lowest total cost and acceptable risk, and for a meaningful slice of any real estate that is no scheduled maintenance at all. An organisation is not mature because it has eliminated run to failure. It is mature because it can tell you, item by item, why each one is on the regime it is on.

3. When run to failure is genuinely the correct answer

This is where most writing on this keyword is dishonest. It treats every instance of reactive work as evidence of failure, which is both wrong and unhelpful, because it leaves practitioners with no vocabulary for the large number of items where scheduling proactive work would be a waste of money. So let me be direct: on a great many items, run to failure is the right engineering answer, and choosing it is a sign of competence rather than neglect.

The logic that legitimises this is not improvised. Reliability-centred maintenance's task-selection reasoning explicitly allows "no scheduled maintenance" as a valid outcome when no proactive task is technically feasible and worth doing. SAE JA1011_202411 sets out the evaluation criteria a process must satisfy before it may legitimately be called RCM, including the required consequence-evaluation and task-selection logic; note that it sets criteria and does not prescribe a process, so it tells you whether your reasoning qualifies rather than handing you a method. SAE JA1012_201108 is the accompanying guide. If you want the method itself in practitioner form, start with the practical introduction to RCM.

Here is the checklist I apply. Run to failure is a candidate only when all of the conditions in the left column hold, and it is disqualified the moment any condition in the right column applies.

Eligibility conditions (all must hold) Disqualifiers (any one rules it out)
The consequence of the failure is low and tolerable in operational terms. Nothing important stops, and nothing downstream is put at risk. The failure has a safety consequence: it could injure or kill someone, directly or through the loss of a protective function.
The item is inexpensive and can be replaced or repaired quickly, with a part that is genuinely held or genuinely available. The failure has an environmental consequence: release, contamination, discharge or breach of an environmental permit condition.
No proactive task addresses the dominant failure mode economically. Either there is no detectable warning to act on, or the task costs more than the failure it avoids. The failure has a statutory or regulatory consequence: an inspection, test or certification is required by law or by the authority having jurisdiction. Statutory work is never a candidate.
The failure is evident. It announces itself to operations under normal circumstances, so somebody knows it has happened. The function is hidden or protective. Nobody finds out it has failed until it is needed, and by then the protection is already absent.
The failure does not damage anything else. The item fails alone rather than taking a shaft, a seal, a coupling or a process batch with it. Failure causes collateral or secondary damage whose cost exceeds the cost of the proactive task you declined.
The decision is documented, owned, and reviewed when duty, criticality or configuration changes. The item is critical and unanalysed. It is on run to failure by default because nobody assessed it. That is not run to failure; it is a reactive posture.

The hidden-function exclusion deserves a sentence of its own because it is the one practitioners most often miss. If a failure is hidden, there is no reactive posture available to you. You cannot react to something you never find out about. A protective device that has quietly failed does not announce itself; it simply is not there when the event it was installed for arrives, and you discover both failures at once. That is precisely why failure-finding tasks exist as a distinct category in RCM task selection: their purpose is not to prevent the failure but to reveal it, on an interval chosen against the probability of the multiple failure. Any item whose function is hidden or protective is therefore outside the run-to-failure decision space entirely, regardless of how cheap the item is.

Deciding which items are eligible for this analysis at all, and in what order, is a criticality question. That framework is in equipment criticality analysis and in the classification scheme in asset criticality classification. Identifying the dominant failure modes to test against the checklist is covered in failure modes and how to analyse them.

Where this checklist does not help you

The checklist is only as good as the failure-mode list you feed it, and most organisations do not have one worth the name. If you cannot state how an item credibly fails, you cannot honestly conclude that no proactive task addresses the failure mode, and a run-to-failure decision made on that basis is a reactive posture with better paperwork. The checklist tests a decision. It cannot manufacture the analysis the decision is supposed to rest on.

4. The symptoms of an organisation stuck in a reactive posture

A reactive posture is diagnosable, and it is diagnosable without a single metric, which is useful because the metrics in a reactive organisation are usually unreliable anyway. These are the signs I look for in the first two days on site.

  • The backlog only ever grows. Not because the backlog is large, which by itself means little, but because its direction is monotonic. Work is added faster than it is closed, quarter after quarter, and nobody can tell you what is in the oldest fifth of it. The measurement discipline that makes a backlog readable rather than frightening is in maintenance backlog and downtime tracking.
  • Preventive work is routinely deferred to cover emergencies. The PM schedule exists and is systematically raided for labour. Compliance is reported against a schedule that was already abandoned in week two of the month.
  • The planner has been absorbed into today's crisis. This is the single most reliable indicator. When the person whose job is next week is spending the day chasing a part for this morning, planning has stopped, and once planning stops the next fortnight is guaranteed to be reactive too.
  • Parts arrive expedited. Courier charges and premium freight become normal rather than exceptional, and the stores function shifts from holding what is needed to procuring what is urgent.
  • Root causes are never examined. Nobody investigates the failure because the next job is already waiting. The item is restored and the record closed with two words. The same item fails again in six weeks and nobody connects the two events because the history is not readable.
  • Shift handovers are entirely about incidents. If the handover never mentions anything proactive, the proactive work is not happening regardless of what the system says.
  • Nobody can tell you why an item is on the regime it is on. Ask about three items at random. In a reactive organisation the answer is "it has always been like that" or nothing at all.

The one symptom that deserves more than a bullet is cultural, and it is the reason the posture persists even when everyone can see it.

Firefighting is visibly rewarded and prevention is invisible. Think about who gets thanked in a reactive organisation. The technician who came in at two in the morning and had the pump back on line by six is a hero, and rightly so, because the effort was real and the outcome mattered. That is a visible event with a name attached. Now think about the engineer who eighteen months earlier changed a lubrication specification so that the pump never failed at two in the morning in the first place. There is no event. There is no call-out. There is nothing to thank anybody for, because the thing that was avoided leaves no trace. Prevention produces an absence, and organisations are structurally poor at noticing absences.

So the incentive gradient points the wrong way. The visible path to recognition runs through emergencies, and the work that reduces emergencies makes the person who does it less visible. Competent planners quietly conclude that firefighting is better for their career than planning, and they are not wrong about their own organisation. Until leadership deliberately makes prevention visible, by reporting avoided failures alongside restored ones, by giving the planner a protected remit that is not raided, and by asking at the review what did not happen this month and why, the culture will keep regenerating the posture faster than any process change can remove it. This is not a soft problem appended to the technical ones. It is the mechanism that holds the technical ones in place.

5. The true costs, honestly and without figures

You will find a widely repeated claim that reactive work costs some multiple of planned work, usually stated as three to five times. I am not going to repeat a number, because I have never seen that figure traced to a methodology that would survive scrutiny, and the multiple depends entirely on the asset, the consequence and the operating context. Treat it as an unsourced industry assertion. The costs below are real and you can identify them in your own operation without borrowing anybody's multiple. The right benchmark is your own baseline, measured before and after you change something.

  • Collateral damage. An item allowed to run to destruction frequently damages what it is attached to. The bearing that seizes scores the shaft. The belt that shreds takes the guard and the pulley. The cost of the failure is not the cost of the item, it is the cost of everything the item touched on its way down, and that cost is almost never captured against the original decision.
  • Secondary failures. The stopped item shifts duty onto something else. The standby unit that has been idle for two years runs continuously for a fortnight and fails in its turn. One reactive event becomes two, and the second one is harder to explain.
  • Safety exposure from unprepared work. This is the cost I weight most heavily. Reactive work is performed under time pressure, frequently out of hours, often without the method statement, the isolation plan, the right tools or the second pair of hands that the same task would have had if it had been planned. Every shortcut that planning exists to prevent is available, and the person taking it is tired.
  • Overtime and expedited logistics. Premium labour rates, call-out allowances, air freight, and paying a supplier for the privilege of being their emergency. None of it buys anything you would not have bought more cheaply with two weeks of notice.
  • Lost output and lost service. Production not made, a building not fit for occupation, a service-level commitment breached. In a contracted facilities environment this is where reactive posture turns directly into a commercial penalty.
  • Data destruction. Reactive work is recorded badly, and badly recorded work destroys the evidence base you would need to fix the problem. The posture erodes the very information required to escape it.

Now the item that matters more than all of the above, because it is the mechanism rather than a line item.

The feedback loop is the central mechanism

Reactive work destroys the schedule. A destroyed schedule means proactive work does not get done. Proactive work not getting done means more items fail unexpectedly. More unexpected failures mean more reactive work, which destroys the schedule again. That is a closed positive-feedback loop, and it explains something the cost list alone cannot: why reactive organisations do not gradually drift back to stability on their own, and why working harder inside the loop makes the loop faster rather than weaker. You do not exit a reinforcing loop by increasing effort. You exit it by breaking one of its links.

Link in the loop What it does to the next link Where you can cut it
Unexpected failure occurs Consumes labour that was committed elsewhere, at short notice Eliminate repeat offenders so the arrival rate falls
Schedule is raided to cover it Proactive tasks slip and then disappear Protect a ring-fenced proactive capacity that cannot be raided
Proactive work is not executed Degradation continues undetected on more items Target the small critical set first rather than the whole register
More items fail unexpectedly Arrival rate of reactive work rises, loop closes Get failure data flowing so you can see which items to attack

6. How to get out, in a sequence that works

The instinct when an organisation recognises the posture is to launch a full preventive maintenance programme. I would advise against it, every time. A large PM programme launched into a reactive environment is consumed by the environment within a quarter: the labour it needs is the labour that is currently firefighting, so the schedule is raided from the first week, compliance collapses, and the programme's visible failure makes the next attempt politically harder. You have to break the loop before you can build the programme. The sequence below is ordered deliberately.

  • Step 1: stabilise by protecting a small proactive capacity. Not a programme. A small, specific, ring-fenced slice of labour, perhaps one technician for part of each week, that cannot be pulled onto today's emergency by anybody below the maintenance manager. It must be small enough to be genuinely defensible and specific enough that raiding it is visible. This single act is what interrupts the loop, because it guarantees that some proactive work happens regardless of the day's noise.
  • Step 2: get the failure data flowing. You cannot target anything without knowing what is failing and how. That means recording the failed item, the failure mode and the consequence on every reactive job, in a form that can be counted later. Keep the coding scheme small enough that tired people at midnight will actually use it; a short list applied consistently is worth far more than a taxonomy applied at random. ISO 14224:2016 is the reference for reliability and maintenance data collection and exchange, originally from the petroleum and gas sector but the most useful published taxonomy model for anyone structuring failure data. A note on a common error: OREDA is a proprietary members-only database, not a standard, and should never be cited as one.
  • Step 3: use criticality to pick the small number of items worth analysing first. Do not analyse the register. Rank by consequence, take the top handful, and analyse those. The point of criticality here is not classification for its own sake, it is permission to ignore most of the estate while you build capability on the part that matters.
  • Step 4: eliminate the repeat offenders. With a few months of readable failure data you will find a short list of items generating a disproportionate share of the reactive work. Take them one at a time and remove the cause rather than the symptom, whether that is a design change, a lubrication change, an operating-practice change or a spares decision. Each one you eliminate permanently reduces the arrival rate of reactive work, which widens the protected capacity from step one without anybody having to approve more labour. This is the step that compounds.
  • Step 5: build planning capability. Only now. Once the arrival rate has fallen far enough that the planner can spend most of the week on next week, planning becomes possible, and once planning is possible a preventive programme can be built that will actually be executed. Attempting this step earlier is what produces the abandoned programmes. The reliability-improvement work that follows naturally from here is covered in equipment reliability and how to improve it.
  • Step 6: make the improvement visible, deliberately. Report what did not fail, not only what was fixed. If the recognition structure keeps pointing at emergencies, the culture will pull the organisation back into the posture no matter how good the technical sequence was.

There is an asset-management frame for all of this if you need to make the case in governance language rather than maintenance language: ISO 55000:2024 for vocabulary, overview and principles, ISO 55001:2024 for the requirements of an asset management system, which is the certifiable one, and ISO 55002:2018 for guidance on applying it, which has not been revised alongside the 2024 pair. Being able to show, item by item, why each one is on the regime it is on is exactly the kind of demonstrable decision basis that framework expects, and it is exactly what a reactive posture cannot produce.

The honest cost of this sequence

This takes quarters, not weeks. Step four is where the arithmetic starts working in your favour and you will not reach it inside a single quarter with readable data behind you. And it fails in one predictable way: the protected capacity from step one is always the first thing surrendered when a bad week arrives. Organisations that abandon this sequence abandon it there, not at a later step, and typically for a reason that was entirely defensible on the day. If the protected capacity is negotiable at supervisor level, the sequence has already failed and the rest of the plan is decoration.

7. What a reactive posture is not

Two clarifications, because the posture gets misdiagnosed in both directions and both misdiagnoses waste money.

It is not the same thing as having a low preventive maintenance budget. A small, well-targeted proactive programme covering the items where proactive work genuinely pays, with everything else on a documented and reviewed run-to-failure decision, is a well-managed estate operating at low cost. It is not a reactive posture. Judging an organisation by the size of its PM programme measures spend, not control, and it drives exactly the wrong behaviour: adding tasks to look proactive.

And the converse, which is the more uncomfortable one: an organisation can be thoroughly reactive while running a very large preventive maintenance programme. If the PM content addresses the wrong failure modes, all that scheduled effort is not preventing the failures that are actually occurring. Compliance reporting looks excellent, labour is fully committed, and items keep failing unexpectedly, because the tasks and the failure modes were never connected to each other. There are estates where a heavy inspection regime is executed faithfully and the dominant failure mode on the critical plant is not addressed by a single task in it. That is a reactive posture with a full schedule, which is the most expensive version of the problem, because it pays for the programme and the emergencies both.

The test is therefore not how much proactive work you do. It is whether the proactive work you do addresses the failure modes that are actually causing your failures. That question is answerable in an afternoon: take your last twenty unplanned events, and for each one ask which scheduled task was supposed to have caught it. The pattern of blanks in that column is your real diagnosis.

The idea to walk away with

Reactive maintenance is a posture, and run to failure is a decision. Everything useful in this subject follows from holding those two apart. If you have analysed an item, concluded that no proactive task is worth doing, recorded that conclusion with an owner and a review trigger, and confirmed that the failure is evident and carries no safety, environmental or statutory consequence, then you are managing that item deliberately and the fact that it will eventually be repaired after it fails is not a criticism of you. If you have not done that analysis, then you are not running it to failure. You are simply waiting, and the failure will decide the timing.

The reason the posture is hard to leave is not that anybody is unaware of it. It is that it is a reinforcing loop with a cultural lock on it: reactive work destroys the schedule, the destroyed schedule generates more reactive work, and the recognition structure rewards the people inside the loop while making the people who would break it invisible. Break the loop at the one link you control, which is the protected proactive capacity, and defend that capacity when defending it is inconvenient.

Final thoughts

If you take one action from this guide, take the smallest one: pick five items at random from your estate and ask why each is on the regime it is on. Not what the regime is, the system will tell you that. Why. The proportion of the five for which somebody can give you an answer that refers to a failure mode and a consequence is the most honest measure of posture I know, and it costs nothing to run.

And be fair to the organisations that come out of that test badly, because most do. Nobody chose this. Reactive posture is what happens by default when an estate grows faster than the analysis behind it. The way out is not more effort, which the loop will absorb. It is a small protected capacity, readable failure data, a short list of items worth analysing, and the patience to let the compounding do its work over several quarters rather than several weeks.

Disclosure

Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.

Stuck in a reactive posture?

Independent advisory on breaking the reactive loop: protected proactive capacity, failure data structure, criticality-led targeting and the run-to-failure decisions worth documenting. 22+ years across utilities, oil and gas, manufacturing, government and facility operations.

Book a conversation

Related reading: Corrective maintenance: a complete guide, Breakdown maintenance: a complete guide, Planned vs unplanned maintenance, Preventive vs predictive vs reactive maintenance, Equipment reliability and how to improve it, Maintenance backlog and downtime tracking. Standards bodies referenced: CEN-CENELEC , SAE International , ISO .

Muhammad Abbas

CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.

Work with me
MAbbaz.com
© MAbbaz.com