mail@mabbaz.com Abu Dhabi, UAE

Breakdown Maintenance · Emergency Response · Reliability

Breakdown Maintenance: A Complete Guide

Every organisation has breakdowns. Very few handle them deliberately. This guide is about the breakdown event itself: what happens from the moment something stops performing its required function to the moment the event is safely closed out and genuinely learned from, and what to put in place beforehand so the response is not improvised under pressure.

Muhammad Abbas September 27, 2026 ~19 min read

Ask a maintenance team how they handle a breakdown and you will get a description of what they do, not of how they decide. Someone notices, someone calls, someone attends, something gets fixed, production resumes. That sequence mostly works, which is why it is rarely examined. But the difference between an organisation that handles breakdowns well and one that merely survives them is not speed. It is whether a handful of specific decisions get made consciously, in the right order, while the pressure to restore output is at its peak.

The message up front: a good breakdown response is not a fast one, it is a controlled one. Make safe before anything else, establish what has actually failed and tell the people who must plan around it, then consciously choose between restoring function quickly and repairing properly now. Preserve the evidence, verify the repair rather than assume it, and close out by recording what was found rather than what was done.

1. What breakdown maintenance means, and the words the trade uses loosely

Breakdown maintenance is the work carried out after an item has stopped performing its required function, in order to restore it. That definition is deliberately narrow: work triggered by a failure that has already occurred, on an item that is no longer doing its job.

The formal vocabulary sits in two places. EN 13306:2017, "Maintenance: Maintenance terminology", is the European standard defining the maintenance types and their relationships; it is a CEN standard with no ISO twin, which surprises people. IEC 60050-192:2015, Part 192 of the International Electrotechnical Vocabulary, covers dependability terms and is the formal home of the failure, fault and downtime vocabulary. Neither is law by itself, and both are paywalled, so check the current published text rather than trusting a definition you read in a blog post, including this one.

Now the honest part. In day to day use the trade treats "breakdown", "emergency" and "reactive" as interchangeable. They are not:

  • Breakdown maintenance is the work itself: the response to an item that has failed. It is an event and a response to that event.
  • Emergency maintenance is a subset distinguished by urgency and consequence, not by cause. A breakdown becomes an emergency when the consequence is safety, environmental, statutory or business critical enough that normal planning is suspended. Most breakdowns are not emergencies, and treating them all as emergencies destroys planning discipline.
  • Reactive maintenance is an operating posture: how far a whole maintenance function is driven by events rather than by plan. For that lens, and run to failure as a strategy choice, see the reactive maintenance guide.
  • Corrective maintenance is the work category breakdown work is recorded under, and it is broader, also covering deferred jobs raised from inspection findings on items that have not yet failed. Those mechanics belong in the corrective maintenance guide.
  • Unplanned maintenance is the scheduling attribute. Breakdown work is usually unplanned, but not always, and some planned work is corrective. The quadrant is set out in planned versus unplanned maintenance.

I am not re-teaching any of those four here. This article stays on the event: what a competent organisation does between the moment something stops and the moment the file is closed.

2. Functional failure against total stoppage: the distinction that gets missed

Here is a commonly missed idea in breakdown management. A failure is a loss of the ability to perform a required function. It is not the same thing as a stoppage.

A chilled water pump specified to deliver a given flow at a given head, now delivering perhaps sixty percent of that because of impeller wear, has functionally failed. It is still running, nothing has tripped, nobody has raised anything. But it can no longer do what is required, and the consequence is absorbed elsewhere: a second pump running that should be on standby, space temperatures drifting, occupant complaints logged against the wrong asset. That is a breakdown, whether or not anyone has called it one. The example is my own and illustrative.

Organisations routinely fail to raise functional failures as breakdowns until the item stops completely, for reasons that are structural rather than careless:

  • The required function was never written down. You cannot detect a failure against an undefined requirement.
  • Nothing alarms. Control systems generally detect stoppage and fault, not degradation against a design duty. A running machine reads as a healthy machine.
  • Redundancy hides it. Duty and standby arrangements absorb exactly this, so well that the failure is invisible until the second unit is also compromised. Loss of redundancy is itself a failure of a required function, and rarely raised as one.
  • Nobody is accountable for output. Operators report what stops. Nobody is asked whether equipment is still meeting duty.
The test worth applying

For each item that matters, can you state in one sentence what it is required to do, in measurable terms, and can somebody see whether it is currently doing that? If not, you can only detect total stoppages, which means a proportion of your breakdowns are running quietly right now and will present later as something worse. Defining required function is unglamorous asset data work, and it is the precondition for everything else here.

The fix is to give people a legitimate route to raise a degraded failure: a work request reason that is neither "broken" nor "routine", and a supervisor who treats a reported loss of duty as a real finding rather than an operator being fussy. This is where equipment reliability work starts paying back.

3. Make safe first, and mean it

Everything else in this guide is negotiable in sequence. This is not. Before diagnosis, before assessment, before anybody phones the production manager, the item and the area around it are made safe: shut down as far as condition allows, hazardous energy isolated and the isolation verified, stored energy dissipated or restrained, and the isolation locked and tagged so it cannot be reversed by somebody who does not know work is in progress.

The point here is timing rather than technique. The moment of highest commercial pressure to restore output is exactly when isolation discipline is most likely to be abandoned. A line is down, somebody senior is asking for a restart time, and the shortcut presents itself as reasonable: just hold the contactor, just try it with the guard off, just have somebody stand by the isolator instead of locking it. The pressure does not change the hazard, only the person's willingness to accept it. An organisation that has decided in advance that isolation is not negotiable under time pressure removes that negotiation from a technician who is in no position to win it.

Competence and jurisdiction, stated once. Hazardous energy isolation must be carried out only by people formally assessed as competent for that equipment and energy type under your own authorisation scheme, and the legal duties differ by country, so nothing here substitutes for your local statutory requirements and safe systems of work. In the United States the federal energy control requirements sit in 29 CFR 1910.147, the control of hazardous energy standard, binding US general industry only, with State Plan states permitted to impose differing rules; in Great Britain the equivalent duties arise under separate general health and safety and work equipment legislation rather than a single lockout regulation. Establish which applies to you first.

One rule worth stating separately: the person who applied the isolation controls its removal. Nobody restores energy on a verbal assurance that the work is finished. I am deliberately not teaching isolation procedure here, because procedure written generically fits nobody's plant. Lock and tag application, group isolation, shift handover, verification methods and the authorisation hierarchy are set out in the lockout tagout safety guide, along with the permit interface where the work needs one. A breakdown response that is safe and nothing else is still not a good response, so read that and move on.

4. Assess and communicate: establish what stopped, and tell the people who must replan

Once the area is safe, establish two things and tell two audiences. Most organisations do the first adequately and the second badly.

What has actually stopped. Not what was reported. Reports arrive as symptoms attributed to the nearest visible asset, and the attribution is wrong often enough that acting on it wastes the first hour. The assessment establishes which item has lost which function, and whether what presented as the failure is the failed item or a downstream consequence. A tripped distribution board is usually reporting a fault on its load, not a fault in itself.

What the operational consequence is. A separate question, answered by operations rather than by the technician. Is output stopped or reduced? Is redundancy lost, so the site is one failure away from a far worse position? Is there a statutory, contractual or environmental implication? Consequence determines urgency, which determines resources. Deciding urgency by how loudly someone complained is the default in most organisations and it consistently misallocates the crew.

Then the communication, where I see the most avoidable damage. Two audiences need different things:

  • Operations and the business need to know what is unavailable, the current assessment, and when the next update comes. They do not need a diagnosis. The most useful discipline is committing to an update time rather than a restoration time. "We do not yet know the cause. Next update at eleven." is a complete and professional message. A guessed restoration time will be wrong, will be planned against, and will cost you credibility you need later in the same event.
  • The maintenance chain needs the technical picture: what is isolated, what has been found, what parts and skills look likely, what help is needed. Escalation should be a defined route, not a matter of who answers the phone.
Where this goes wrong

Silence during a breakdown is filled by speculation. The absence of updates turns a technical problem into an organisational incident, because people who cannot get information start making their own arrangements and arriving on site to see for themselves. A nominated single point of communication, updating on a committed cadence, is worth more to a difficult breakdown than an extra technician.

5. Decide the objective: restore function fast, or repair properly now

This is the section most breakdown content omits, and the most valuable decision in the sequence. Once you know what has failed and what the consequence is, and before anyone starts work, somebody chooses the objective of this intervention. There are two legitimate objectives:

  • Restore function fast. Get the required function back, accepting a temporary or partial solution: the reconditioned spare rather than a four week wait, the failed component bypassed and run manually, a cross connection from the adjacent system, a hired unit. The function returns; the defect remains.
  • Repair properly now. Take the outage you are already in and do the correct, permanent repair, accepting a longer restoration in exchange for not coming back.

Both are defensible. Which is right depends on the consequence of continued loss of function, the availability of correct parts and skills, whether the temporary solution introduces new risk, and when the next planned opportunity realistically falls. Restoring fast on a critical utility asset with a long parts lead time is obviously correct. Rushing a permanent repair because somebody wants the item back forty minutes sooner is obviously not.

The failure is not choosing either one. It is making the choice by accident, under pressure, without anyone naming it, and then never returning. That is how organisations accumulate a hidden population of temporarily restored equipment: the jumper still fitted, the valve still chained open, the failed sensor still bypassed in the control logic, all forgotten because the temporary fix worked and the job was closed as complete.

The rule that prevents most of this

If you restore temporarily, the permanent repair is raised as a separate job at the same moment, before the original job is closed, with a review date and a named owner. Not later, not at the end of the shift, not in the handover notes. A temporary restoration with no corresponding permanent job on the system is not a repair, it is a liability you have chosen to forget about, and whoever inherits it will not know it is there.

There is also an authority question, and it belongs to preparation rather than to the event. Who may make this call, at each level of consequence, at three in the morning? Unagreed, it defaults to whoever is most senior on the phone, usually somebody optimising for restoration time because that is the only variable they see.

6. Preserve the evidence before it disappears

A breakdown is the only opportunity you will get to understand that particular failure, and the evidence has a very short life. Within an hour of the repair starting, the failed component has been binned, the spilled oil mopped, the burnt terminal cleaned up and re-terminated, the operating conditions overwritten in the historian's buffer, and the operator who saw it happen has gone home. What remains is a work order reading "replaced pump", and nobody will ever be able to say why it failed.

This loss is one of the largest reasons the same breakdown recurs. Not poor repair work: nobody can establish the cause, so nobody addresses it, so the mechanism stays in place and produces the same failure again. Capturing evidence costs a few minutes. What to capture:

  • The failed part itself. Labelled with the asset, date and work order reference, and quarantined rather than binned. A failed bearing tells a competent examiner what killed it; a photograph of one usually does not. This discipline requires nothing but a shelf and a rule.
  • Photographs before disturbance. As found condition, the installation as it sat, the failure surface, collateral damage, local gauges and the nameplate. Cleaning destroys the evidence that explains the failure.
  • The operating conditions at the time. Load, duty, speed, temperatures, pressures, running hours, recent starts and stops. Much of this sits in a BMS, SCADA or historian and will be aggregated away within days, so export it now rather than requesting it next month.
  • What the operator observed. What changed first, what it sounded or smelled like, whether there were prior symptoms. Operators very often know an item had been behaving oddly for weeks. Nobody asks.
  • Recent history of the item. What was done to it last, by whom, and whether the failure followed an intervention. Maintenance induced failure is common and systematically under recorded, because the person best placed to spot it is the one least motivated to.

What you do with the evidence afterwards is a separate discipline, covered in failure analysis methods, process and examples. Standardised failure data capture is addressed by ISO 14224:2016, the third edition of the reliability and maintenance data collection standard from the petroleum, petrochemical and natural gas industries; its taxonomy is genuinely useful well outside them. Worth noting, because the two get conflated constantly: OREDA is a proprietary members only database, not a standard. ISO 14224 is the standard that grew out of that work.

7. Diagnose, repair, verify, return to service

With the area safe, the objective chosen and the evidence captured, the technical work proceeds. Two points, rather than a repair tutorial.

Diagnosis is not the same as identifying the broken part. Finding the failed bearing is identification. Diagnosis asks why it failed, and the answer is frequently not the bearing: misalignment, over or under greasing, contamination past a failed seal, the wrong bearing fitted last time, a duty the equipment was never selected for. Replace the part without forming at least a provisional view on the mechanism and you have scheduled the next breakdown rather than prevented it.

Verification is a distinct step, and "it started" is not verification. This is the step most often collapsed into the repair: the equipment runs, the noise has gone, everyone disperses. Starting proves only that the item can start. Verification proves the required function has been restored, checked against the duty defined back in section two: is the flow, pressure, temperature or output back at requirement, under representative load, for long enough to be meaningful? It also confirms the protective functions the repair touched are working, that temporary measures and isolations have been formally removed, that guards and interlocks are reinstated and proven, and that instrumentation is not still reading a value somebody forced during testing.

Return to service is then a deliberate handover: operations formally accepts the item back, knowing whether it was permanently repaired or temporarily restored, what limitation remains, and what to watch for. A verbal "it is running again" is how a temporarily restored machine ends up operated as though fully repaired.

The uncomfortable one

A proportion of breakdowns are followed within days by a second breakdown on the same item. Some are genuinely a different failure. Many are the first repair not having addressed the mechanism, or the verification step having been skipped so a partially effective repair was accepted as complete. If you track nothing else about your breakdown performance, track repeat failures on the same item within a short window. It is the most honest signal available about the quality of your response, and it is uncomfortable to look at, which is why almost nobody does.

8. Close out and learn: record what was found, not what was done

The close out is where the event either becomes organisational knowledge or evaporates, and the distinction that matters is between recording the activity and recording the finding. "Replaced motor" is activity, and tells a future reader nothing. "Motor failed on winding insulation breakdown, drive end bearing also seized, water ingress at the terminal box gland, gland found loose and refitted with new seal, motor exchanged" is a finding: one sentence longer, and the difference between a history you can analyse and one you cannot.

A close out worth the name captures:

  • What was found as distinct from what was done, including the provisional mechanism even where uncertain. "Cause not established" is a legitimate record. A blank field is not.
  • Failure coding against a controlled structure rather than free text, so the event is countable. Coded data reveals that the same failure mode has occurred repeatedly across a class of equipment, a pattern no amount of reading individual work orders will surface. The structure is set out in failure codes: problem, cause, action.
  • Parts actually consumed. Unrecorded consumption is how a critical spare quietly reaches zero.
  • Downtime and labour, against a definition you have written down. The industry genuinely disagrees about what clock starts and stops: active hands on repair only, or total downtime including waiting for a decision, parts and access. IEC 60050-192:2015 provides the dependability vocabulary, but the disagreement is about local application rather than the words, and it is the source of most arguments about repair time metrics. Pick one definition, apply it consistently, and state which you used. See MTTR meaning, formula and how to improve it and maintenance backlog and downtime tracking.
  • Any follow up raised, particularly the permanent repair behind a temporary restoration, with its review date and owner.

Then the trigger for deeper analysis. Not every breakdown warrants a formal investigation, and pretending otherwise guarantees none get one properly. Define the triggers in advance and apply them mechanically: any actual or credible potential safety or environmental consequence, any event above an agreed cost threshold, any repeat of the same failure mode on the same item within an agreed window, any failure of a protective function, and any failure on a critical item. Everything else gets a good close out record and feeds the data.

Where a formal investigation is triggered, there is an international standard for it, which surprises people: IEC 62740:2015, "Root cause analysis (RCA)", adopted in Europe as EN 62740:2015. It sets out principles and process steps and describes recognised techniques including the "Why" method commonly called five whys, the Ishikawa or fishbone diagram, event and causal factor charting, causes trees and fault trees. It covers analysis after the event only and excludes assigning blame. So the correct thing to say about five whys is not that it is unstandardised, but that no standard prescribes how to run it while IEC 62740 recognises it as one of several techniques. Method selection is covered in root cause analysis methods and step by step guide.

9. The response sequence at a glance

The whole sequence, with the failure mode of each step and what should end up on the record. The value is in the third and fourth columns rather than the first.

Step Objective What typically goes wrong What to record
1. Make safe Remove the hazard before anyone touches the equipment Isolation shortcuts taken under pressure to restore output; energy source missed; isolation removed by someone other than the applier Isolation applied and verified, by whom, under which authorisation; permit reference if applicable
2. Assess Establish which item lost which function, and the operational consequence Acting on the reported symptom rather than the actual failed item; urgency set by who complained loudest Item and function lost; as found condition; consequence and urgency classification
3. Communicate Tell operations what is unavailable and when the next update lands Silence, then a guessed restoration time that is planned against and missed Notification time, who was told, commitments given, update times
4. Decide the objective Consciously choose restore fast or repair properly now The choice is made by accident under pressure and never revisited; no named decision authority Objective chosen, who decided, rationale, and the permanent job raised if temporary
5. Preserve evidence Capture what is needed to understand the failure later Failed part binned, scene cleaned, conditions rolled off the historian, operator gone home Part retained and labelled; photographs before disturbance; operating data exported; operator account
6. Diagnose Form a view on the mechanism, not just the broken part Part identified, mechanism never considered, so the failure is scheduled to repeat Provisional mechanism, or an explicit note that it could not be established
7. Repair Execute the repair consistent with the chosen objective Scope drifts beyond or below the agreed objective without anyone deciding to change it Work performed, parts consumed, labour, any deviation from specification
8. Verify Prove the required function is restored under representative load "It started" accepted as proof; protective functions and interlocks not proven; forced values left in place Function checked against duty, values measured, protective functions proven, temporary measures removed
9. Return to service Formal handover back to operations with any limitation stated Informal handover, so a temporarily restored item is operated as fully repaired Acceptance by whom and when; remaining limitations; what to watch
10. Close out and learn Turn the event into countable, usable history Activity recorded instead of finding; codes left blank or applied inconsistently; no analysis trigger applied Findings, failure coding, downtime on an agreed definition, follow ups, analysis trigger decision

10. Preparing before the breakdown so the response is not improvised

Most of what makes a breakdown response slow, unsafe or wasteful is decided long before the breakdown. None of the items below needs a budget case, a project or a piece of software. They need somebody to do the work once, deliberately, for the equipment that matters.

Preparation item What good looks like What it costs you if it is missing
Required function defined For every critical item, a one line measurable statement of what it must deliver, held with the asset record You can only detect total stoppages, so functional failures run on undetected until they become worse failures
Critical spares held deliberately Holdings chosen from failure consequence and lead time, not from what was left over; location known; stock accurate; insurance spares identified as such Restoration time becomes a procurement lead time, and the restore fast option disappears entirely
Isolation points known and labelled Every energy source identified per item, physically labelled, and cross referenced to the asset; isolation points confirmed as lockable Time lost hunting for isolators, and the real risk of an unidentified energy source being missed under pressure
Drawings and manuals findable Current schematics, single line diagrams, control descriptions, O&M manuals and settings, retrievable in minutes by the person on shift Diagnosis by trial and error; parameters guessed; the site relies on one individual's memory
Contact and escalation routes agreed Named roles with out of hours contacts, specialist contractors with response terms already in place, OEM support arrangements confirmed live The first hour is spent on the phone, and the specialist you need has no contract and no obligation to attend
Decision authority named Written authority, by consequence level, for the restore fast versus repair properly call, and for authorising overtime, hire and expedited purchase, valid at three in the morning The call defaults to whoever is most senior on the phone, optimised for restoration time alone
Competence and authorisation current Assessed and recorded authorisation to isolate and to work on each equipment and energy type, with enough coverage on every shift Either the work waits for an authorised person, or an unauthorised person does it anyway
Evidence capture arranged A labelled quarantine shelf for failed parts, an agreed photograph list, and a known route to export operating data before it rolls off The same breakdown recurs indefinitely because nobody can ever establish why it happened
Failure coding structure in place A short, controlled, genuinely used code set that technicians can apply in seconds without guessing History that cannot be counted, so patterns across an equipment class stay invisible
Downtime definition written down One agreed definition of what the clock includes, published and applied consistently, stated whenever the number is reported Metrics that cannot be compared between months, sites or contractors, and arguments instead of decisions

On spares, the question is not what fails often but what combination of consequence and lead time makes waiting unacceptable, which is why a cheap long lead item on a critical duty beats an expensive one available next day. Software has a modest role here. Any competent maintenance management system holds the asset record, the isolation references, the document links, the coded failure history and the follow up job, which is useful because the information is then in one place when somebody needs it. The system does none of the thinking. An organisation with a well configured platform and no agreed decision authority still makes the restore fast call badly, just with a better audit trail of having done so.

11. The honest limits of getting good at breakdowns

Everything above will make your breakdown response better. None of it will make your plant more reliable, and the two get confused constantly. Breakdown response cannot be optimised into a maintenance strategy: it is what you do when your strategy has already been overtaken by events. Compress the response, sharpen the decisions, tighten the safety and improve the records, and you will still be responding to the same number of failures. The aim is fewer breakdowns and less consequential ones, not faster heroics.

The thing nobody wants to hear

An organisation that becomes genuinely excellent at breakdown response can mask a serious reliability problem indefinitely. The crew is fast, the spares are right, the outages are short, and the equipment is quietly failing at the same rate it always did. Because the consequence never reaches the business, nobody funds the reliability work, and the capable response team becomes the reason the underlying problem is never addressed. If your breakdown response is excellent, that is exactly when to ask how many breakdowns you are responding to, and whether that number is moving.

There is a cultural version of the same trap: breakdown response is visible, urgent and rewarded, while nobody is congratulated for the failure that did not happen. The counterweight is to report the breakdown count and the repeat failure count with at least the same prominence as the response time, so the organisation can see whether it is getting better or merely getting faster.

The other limit is honesty about metrics. You will be asked what a good response or repair time looks like, and you should resist answering with a number from elsewhere. There is no credible universal benchmark, because it depends entirely on your equipment, consequences, spares position, labour model and definitions. The only defensible target is one set against your own measured baseline.

The idea to walk away with

A breakdown is not primarily a repair problem. It is a sequence of decisions taken under pressure, and the outcome is set by whether those decisions were made deliberately. Make safe first and without negotiation. Establish what has actually lost function and tell the people who must replan around it. Choose consciously between restoring fast and repairing properly, and if you restore temporarily, raise the permanent job in the same breath. Preserve the evidence. Verify rather than assume. Record what you found, not what you did.

Get those right and the event is handled well. But remember what the event is: a signal that something in the strategy did not hold. The organisations genuinely good at this are not the ones with the fastest response, but the ones whose breakdown count is falling.

Final thoughts

If you want one practical place to start, start with the temporary repairs. Find out how many items on your site are running on a temporary restoration with no permanent job raised against them. In most organisations nobody can answer that, which is itself the answer. Fixing that one discipline, that a temporary restoration always creates a permanent job with an owner and a review date, changes more than any process document you could write.

After that, define required function for your critical equipment, so your people can raise a failure before the machine stops. The examples here are my own and illustrative, but the pattern is consistent: the expensive breakdowns are rarely the ones nobody saw coming. They are the ones somebody noticed, in a degraded form, weeks earlier, and had no legitimate way to report.

Disclosure

Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.

Reviewing how your site handles breakdowns?

Independent advisory on breakdown and emergency response design, failure data capture, temporary repair control and the reliability reporting that shows whether the underlying picture is improving. 22+ years across utilities, oil and gas, manufacturing, government and facility operations.

Book a conversation

Related reading: Reactive maintenance: a complete guide, Corrective maintenance: a complete guide, Planned vs unplanned maintenance, Lockout tagout safety guide, Failure analysis methods, Root cause analysis methods, MTTR: meaning and formula, Equipment reliability, Failure codes: problem, cause, action.

Primary sources: CEN-CENELEC (EN 13306), IEC (IEC 60050-192, IEC 62740), ISO (ISO 14224), US OSHA (29 CFR 1910.147, US federal only). Standards referenced here are paywalled and revised; check the current published text before relying on any definition.

Muhammad Abbas

CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.

Work with me
MAbbaz.com
© MAbbaz.com