mail@mabbaz.com Abu Dhabi, UAE

Reliability Engineering · Maintenance Programmes · Asset Management

Equipment Reliability and How to Improve It

Most equipment reliability programmes start in the maintenance department, which is the last place reliability is actually decided. This is a practitioner's guide to what reliability means as a property, where it is really determined, which improvement levers earn their keep, how to sequence a programme depending on how good your data is, and where reliability work is the wrong answer altogether.

Muhammad Abbas September 27, 2026 ~22 min read

Ask a maintenance manager how they intend to improve equipment reliability and you will usually get an answer about maintenance: more preventive tasks, tighter schedule compliance, a condition monitoring system. Ask where the reliability of that equipment was actually determined and the answer is almost always somewhere else, in a specification written by someone who has left, an installation signed off under schedule pressure, and an operating regime nobody has compared against the design envelope in years. That gap is the biggest reason improvement programmes underperform.

The message up front: maintenance is the fourth lever on equipment reliability, not the first. Design and selection, installation and commissioning quality, and operating practice within the design envelope all sit ahead of it, and all three are usually unexamined. Start with the failure data, allocate effort by criticality, get the basics of precision maintenance right, and treat condition monitoring as a later step rather than a first purchase.

1. What equipment reliability actually means as a property

Reliability is a precise engineering property, and the precision is worth recovering because the loose version of the word makes most reliability claims impossible to argue with. Formally, reliability is the probability that an item performs its required function, under stated conditions, for a stated period of time. Four elements, all load bearing. The formal vocabulary sits in IEC 60050-192:2015, Part 192 of the International Electrotechnical Vocabulary, covering dependability terminology and superseding IEC 60050-191:1990. The European maintenance terminology counterpart is EN 13306:2017 (a CEN standard with no ISO twin), which defines the maintenance types used below. Both are voluntary and paywalled, and neither is law anywhere by itself.

The two halves people drop are stated conditions and stated period. Reliability gets quoted as a bare property of the asset, as though it travelled with the nameplate. It does not. A pump that is reliable moving clean water at its design duty point is a different pump when it is throttled hard, run at part load on a drive it was never sized for, or handling a fluid with more solids than the specification assumed. Drop the conditions and the period and a reliability claim becomes unfalsifiable. Nobody can test "this equipment is reliable". Everybody can test "this equipment will perform its duty at rated flow and head, on the specified fluid, in the specified ambient, for the specified interval between overhauls". The first is marketing; the second you can hold a supplier to and later disprove from your own history. Reliability also needs separating from its neighbours, availability and maintainability, which have different levers: see reliability vs availability vs maintainability and the reliability engineering pillar. This article publishes no formulas and no target figures; the measurement side belongs to reliability metrics: MTBF, MTTR and availability and MTBF explained.

2. Where reliability is actually determined

Reliability is set in four stages, in a strict order, and the leverage decreases sharply as you move down. Understanding the order matters, because improvement programmes commonly begin at stage four.

  • Stage one: design and selection. Inherent reliability is fixed before the asset arrives. Material choices, bearing arrangements, sealing design, duty margin, redundancy, protection philosophy and the honesty of the sizing calculation belong here, as does selection: the right machine specified for the wrong duty is a design failure. Whatever is set here is a ceiling.
  • Stage two: installation and commissioning quality. The most under-rated stage, and the one that quietly destroys the most reliability. Alignment, foundation and grouting quality, pipe strain on machine casings, cable terminations, cleanliness at assembly, the correct lubricant rather than whatever was on the shelf, protection settings verified, commissioning tests performed rather than signed. An asset can leave the factory capable and reach steady-state operation already compromised.
  • Stage three: operating practice within the design envelope. Start and stop frequency, whether the asset spends its life near a design point or well away from it, whether trips are reset repeatedly without investigation, whether operators report abnormal noise, heat, leakage or vibration. Operations decides far more of the achieved reliability than the organisational chart suggests.
  • Stage four: maintenance. It cannot exceed the inherent reliability set at stage one. What it can do is stop achieved reliability falling below inherent reliability, and detect and eliminate degradation introduced at stages two and three. A real job, but not the same job as creating reliability.

The honest consequence is worth saying plainly to any sponsor: a badly specified or badly installed asset cannot be maintained into reliability. You can make its failures less surprising and plan around it. You cannot maintain away a duty mismatch, an undersized bearing arrangement, a chronic misalignment locked in by pipe strain, or a machine operating permanently outside the envelope it was sold for. If a category of assets consumes effort endlessly without improving, look upstream before adding tasks. This is also why programmes living entirely inside maintenance plateau: the first three stages sit with projects, procurement, engineering and operations.

3. Defect elimination as a practice

Defect elimination changes reliability directly rather than managing its consequences: find the cause of a failure and remove it, rather than restoring the asset and waiting for the same failure again. The difficulty is organisational, because it requires treating repeat work as a signal rather than as normal. In many operations the same coupling is replaced quarterly and the same seal fails every few months, and every one of those jobs is executed competently and counted as productive work. The history looks healthy; the behaviour is pathological, because the organisation is being paid to absorb a defect rather than remove it. The tell: a job that feels routine to the technician who has done it many times is a defect the organisation has learned to live with. A workable practice needs four things:

  • A trigger that is not discretionary. Repeat failures on the same asset within a defined window, any failure above a consequence threshold, and any failure on a top-criticality asset should enter the queue automatically. If entry depends on somebody feeling motivated, the queue stays empty during busy periods, which is to say always.
  • Analysis proportionate to consequence. Most defects need a short structured conversation with the people who did the work and a look at the history; a small number need a formal investigation. A real international standard exists here: IEC 62740:2015 covers root cause analysis, with an annex describing recognised techniques including the Why method and the Ishikawa diagram, covering only after-the-event analysis and excluding blame. So no standard prescribes how to run 5 Whys, but IEC 62740:2015 recognises it as one of several techniques. See root cause analysis methods.
  • An action owner outside maintenance where the cause is. A specification error belongs to engineering; an operating practice belongs to operations. Closing such an action with a maintenance task is how defects become permanent.
  • Verification that the defect stopped. Not that the action was completed: that the failure stopped recurring. This step is rarely done, which is why so many closed actions coexist with unchanged failure rates.

Understanding which failure modes you are eliminating is a prerequisite: see failure modes and how to analyse them and the bathtub curve explained.

4. Precision maintenance: the highest return and least glamorous lever

If I could change one thing in a typical maintenance organisation, it would not be a system or a monitoring technology. It would be the precision of the work already being done. Precision maintenance means performing existing tasks to a defined engineering standard rather than to the standard of good enough to run, and it is the closest thing to a free reliability improvement that exists, because the work is already being paid for. The elements are unremarkable, which is why they get skipped:

  • Alignment. Shaft alignment to a defined tolerance, measured and recorded, soft foot checked, pipe strain considered. Aligned by eye, or aligned at installation and never re-checked after a pipework modification, is not aligned.
  • Balancing. Rotating assemblies balanced to a specified grade after repair, not just reassembled. An unbalanced rotor consumes bearing and seal life quietly.
  • Torque. Fasteners tightened to specification, in sequence, with a calibrated tool. Foundation bolts, bearing housings, flanges and electrical terminations all fail characteristically when this is left to judgement.
  • Cleanliness. Contamination control wherever a system is opened: hydraulic and lubrication systems, bearing housings, electrical enclosures.
  • Lubrication practice. Correct product, quantity, interval and application method, clean transfer equipment, labelled and segregated storage. Over-greasing a bearing damages it, mixing incompatible greases damages it, a contaminated grease gun damages it.

It is under-invested because it is invisible when it succeeds. Nobody presents a slide about torque wrenches. But the failures it prevents are the same failures condition monitoring programmes spend real money detecting, and it is a strange sequence to buy detection for failures you are still manufacturing. Making it real means specifying tolerances in the job plan, providing and calibrating the tools, training the practice, and recording as-left values so the next person knows what they inherited. That last point converts precision maintenance from an individual skill into an organisational property.

5. Operator involvement and basic care

Operators are in front of the equipment continuously and maintainers are not. That asymmetry works in two directions: operators detect early symptoms, a new noise, a warm bearing housing, a weeping seal, a changed vibration felt through a handrail, long before any monthly route would, and operators cause or prevent a large share of the damage through how they start, load, stop and respond to the machine.

Basic care is the practical form: cleaning that also serves as inspection, simple checks of level and condition, minor lubrication where competence and access permit, and above all a low-friction way to report an abnormality and see something happen as a result. That last clause is where these schemes die. Raise three observations, hear nothing, and you have been trained not to report. Two cautions: do not use operator care as cost reduction disguised as engagement, because transferring tasks without training, tools and time produces worse work rather than cheaper work, and be deliberate about scope, since cleaning, inspecting and reporting suit operator ownership almost everywhere while intrusive tasks on hazardous equipment generally do not.

6. The failure-data foundation

You cannot improve what you cannot see, and in most organisations reliability improvement stalls for lack of visibility rather than lack of effort. Failure history is the instrument panel for the whole programme. If it is unreliable, every downstream decision, which assets are worst, which failure modes dominate, whether an intervention worked, is guesswork with a spreadsheet attached.

The relevant international standard here is central rather than decorative. ISO 14224:2016, third edition, covers the collection and exchange of reliability and maintenance data for equipment, written for the petroleum, petrochemical and natural gas industries but widely used outside them. It gives you a taxonomy: a consistent way to classify equipment, describe failures and their modes, and structure the data so records collected in one place can be compared with records collected elsewhere. That comparability is what ad hoc in-house coding schemes destroy. One clarification, because it is misstated constantly: OREDA is a proprietary members-only database, not a standard. ISO 14224 is the taxonomy and format standard that grew out of that work. The foundation has three components, all usually needing work:

  • Failure coding. A structured record of what was wrong, why, and what was done, applied consistently. Free text is not data, and a field called "fault" with two hundred spellings is not data either. The structure that works is covered in failure codes: problem, cause, action; the point here is that reliability improvement is downstream of it.
  • Downtime capture. When the asset stopped performing its function, when it resumed, and how that interval breaks down between detection, response, waiting for parts, active repair and return to service. Without the breakdown you cannot tell a reliability problem from a logistics problem.
  • History quality. Enough narrative and evidence that somebody two years later can understand what happened: as-found condition, as-left values, parts consumed, what was ruled out. Technicians resist this, reasonably, because it is usually demanded without explaining what it is for.

Presented as compliance, data capture is completed badly; presented as instrumentation, with results visibly fed back to the people who captured them, quality improves markedly.

An illustrative case, my own construction

Consider a site with a large pump fleet where the failure cause field is blank on most records. The team cannot say whether seals, bearings or couplings dominate, so it commissions fleet-wide vibration monitoring. Months later it has trend charts and still cannot say what is failing or why, because trends describe symptoms while the absent cause data described the defects. This is my own illustration, not a specific engagement, but the pattern recurs.

7. Criticality-led allocation of effort

A programme spread evenly across an asset register has declined to make decisions. Assets differ enormously in what their failure costs, in safety, environmental, production, service and regulatory terms, and effort should follow that difference. Criticality assessment converts the principle into an allocation: a defensible ranking, which gives you permission to do less at the bottom of the register, and that permission is the real deliverable. Most organisations know intuitively which assets matter, but without a documented ranking they cannot justify reducing attention on the rest, so effort stays spread thin. The methodology is covered in asset criticality classification and in more depth in the equipment criticality analysis guide.

Two practical points. Keep the scale coarse: three or four bands are enough, since a ten point scale invites argument about a six versus a seven while nothing gets decided. And criticality is a property of the asset in its context, not of the asset type: the same model of chiller is a different criticality serving a data centre than a staff canteen, and ranking by asset class silently reintroduces the even spread you were escaping.

Where asset management governance is in scope, the relevant family is ISO 55000:2024 (vocabulary and principles), ISO 55001:2024 (requirements, the certifiable one) and ISO 55002:2018 (guidance, not revised alongside the 2024 pair). Certification covers the management system rather than reliability outcomes.

8. Condition monitoring where the failure mode develops detectably

Condition monitoring is a genuine lever with one strict qualifier: it only works on failure modes that develop detectably over an interval long enough to act on. Bearing degradation, insulation deterioration, lubricant contamination and progressive erosion all announce themselves. Sudden fractures, control board failures and many electronic faults do not, and no sensing predicts them because there is no developing signature to sense. So: establish the dominant failure modes, ask which develop detectably, ask what technique detects that development, and ask whether your measurement interval is comfortably shorter than the development interval. A technique deployed without that reasoning produces data and no decisions. The techniques are covered in condition monitoring techniques.

On standards, the programme-level wrapper is ISO 17359:2018, with ISO 13379-1:2025 for diagnostics and ISO 13381-1:2025 for prognostics. One point misrepresented in tenders: the ISO 18436 series certifies people, not organisations (ISO 18436-2:2014 vibration, ISO 18436-7:2014 thermography). A company cannot be ISO 18436 certified; individuals can, so if you are buying analysis rather than raw data, ask about individual qualification.

9. PM task content review: compliance is not effectiveness

The observation that changes how people look at their preventive maintenance programme: PM compliance can be excellent while PM effectiveness is nil. Compliance measures whether scheduled tasks were completed on time; effectiveness asks whether those tasks address failure modes that actually occur on that asset. Those are independent properties. Tasks inherited from a vendor manual or a previous contractor can be executed with perfect discipline and still do nothing for reliability, because they were never matched to this asset's failure modes in this application. Compliance is easy to measure and effectiveness is not, so organisations manage what is easy and wonder why reliability has not moved.

The measurement side is covered in preventive maintenance KPIs and schedule compliance, and programme design in the preventive maintenance complete guide. What belongs here is the content review: for each task, what failure mode does it address, does that mode occur on this asset, does the task detect or prevent it, and is the interval related to anything other than habit. Such a review must be willing to delete: inspections of components that never fail, intrusive work that adds more risk than it removes, duplicated checks across overlapping routines.

Where a formal method is wanted for setting policy per failure mode, reliability centred maintenance is the structured route and SAE JA1011_202411 is the relevant document: it sets the evaluation criteria defining what may legitimately be called RCM, and does not prescribe a process, so a process failing any criterion is not RCM whatever it is marketed as. The companion guide is SAE JA1012_201108. RCM is covered in the RCM introduction; here it is one lever among several, and a heavy one.

10. Design and specification feedback into procurement

This is the lever with the longest payback and the highest ceiling, and it is rarely pulled. Every failure is information about what to specify differently next time, and in most organisations it never reaches the people writing the next specification, so the same weakness is purchased repeatedly while maintenance absorbs the consequences. Making the loop exist is procedural: a route by which recurring failure findings reach engineering and procurement in usable form, a maintained set of lessons attached to specifications and tender documents, reliability and maintainability requirements written into purchase specifications, and acceptance criteria at commissioning that test what was specified. Adding a requirement to a tender is easy; verifying it at handover makes it real. Be realistic about the timescale: this improves assets you have not bought yet, which is why it loses every prioritisation contest and must be protected deliberately.

11. Spares and logistics: important, but not reliability

Spares availability, storeroom accuracy, repairable rotation and supplier lead times all matter, and are routinely presented as reliability improvements. They are not. They affect how long a failure keeps the asset out of service, which is availability and maintainability, not the probability of the failure happening. The conflation distorts priorities: a spares project improves downtime figures, gets credited as reliability progress, and the underlying failure rate stays unexamined. Spares work should be justified on downtime and cost; reliability work on failure frequency and consequence. The separation is covered in reliability vs availability vs maintainability.

12. The levers side by side

The final column matters most, because the commonest programme failure is not choosing a bad lever but expecting a lever to fix something it structurally cannot.

Lever What it addresses Typical effort Where it fits in sequence What it will NOT fix
Failure data foundation Visibility of what fails and why Moderate, mostly behavioural and sustained First, always Nothing on its own; it enables, it does not improve
Criticality ranking Where to spend effort at all Low, a workshop plus discipline to maintain Early, alongside the data work Any failure mode; it allocates, it does not intervene
Precision maintenance Damage introduced by the maintenance itself Moderate: tolerances, tools, training, recording Early, highest return per unit effort Design inadequacy or out of envelope operation
Defect elimination Recurring failures with removable causes Moderate and continuous; needs an owner Early, once data supports it Random failures with no addressable cause
Operator care Early detection and operating induced damage Low to moderate; heavy on sustained engagement Early, cheap, quick to show value Anything needing specialist diagnosis or intrusion
PM task content review Tasks addressing modes that do not occur Moderate; engineering judgement intensive Middle, once failure modes are known Poor execution quality of the tasks that remain
Condition monitoring Detectably developing failure modes High: hardware, competence, interpretation, routine Later, after data and precision basics Sudden failures, and any defect you keep creating
Design and spec feedback Inherent reliability of future assets Low effort, high organisational difficulty Structural, start early, benefits arrive late Anything already installed
Spares and logistics Downtime duration, not failure frequency Moderate to high, mostly capital and process Parallel track, justified separately Reliability itself; it is an availability lever

13. How to sequence a programme

Sequence matters more than lever selection, and the right sequence depends on the quality of your starting data. Two organisations with identical registers and budgets should do different things first if one can see its failures and the other cannot.

Dimension Data-poor starting point Data-mature starting point
What you can see Work orders exist; causes blank or inconsistent; downtime not captured or captured as a single duration Failures coded to a consistent taxonomy; downtime broken down; history readable after the fact
First move Fix coding and downtime capture on the top criticality band only. Do not attempt the whole register. Pareto the failures, pick the dominant modes, open defect elimination on them
Second move Precision maintenance basics, which need no data to justify and reduce self inflicted failures immediately PM task content review against the known modes; delete as well as add
Third move Coarse criticality ranking, three or four bands, to stop effort spreading evenly Targeted condition monitoring where modes develop detectably and intervals permit
Defer Condition monitoring purchases, formal RCM, reliability modelling. All consume data you do not have. Fleet wide monitoring. Stay selective even when the data would support more.
Evidence of progress Coding completeness and usefulness of the history, not yet reliability outcomes Recurrence of specific eliminated defects, measured against their own prior history
Main risk Trying to fix the whole register at once and abandoning the effort Analysis becoming the deliverable instead of the intervention

On whether to buy a condition monitoring system first: almost always no. It presumes you know which failure modes matter, which in a data-poor organisation you do not. It presumes the failures it detects are not ones you are still manufacturing through imprecise work. And it front-loads capital and licence cost before the programme has produced evidence, making the second year of funding a harder conversation. Monitoring is a good lever and a poor first lever.

14. The organisational reality: ownership and protected time

Everything above assumes somebody is doing it, and this is where reliability programmes most reliably die. A programme without a named owner and protected time reverts to firefighting inside a quarter, because that is roughly how long a normal run of breakdowns, absences and unplanned projects takes to consume unprotected capacity. Reactive work has a deadline and a visible consequence for missing it, reliability work has neither, and when the same person holds both, reactive wins for entirely rational reasons. The person is not weak, the structure is.

So the structural answers are the ones that work: a named individual whose primary accountability is reliability rather than restoration, time genuinely ring-fenced rather than nominally allocated, a small standing forum with authority to assign actions outside maintenance, and a sponsor who asks about defect recurrence rather than only about backlog. One warning about that sponsor: reliability improvement shows its clearest evidence as the absence of failures, a harder story than a completed project, so a sponsor needing quarterly visible wins will unintentionally redirect the programme toward whatever is most visible. Set expectations about the shape and timing of the evidence at the start.

15. The honest limits of reliability improvement

Three limitations deserve stating plainly, because a programme sold without them gets judged against expectations it was never going to meet.

It is slow, and attribution is genuinely hard. Defect elimination changes failure rates over months to years, and by then the operating regime, the workforce, the load and the product mix have all changed too. Proving your intervention caused the improvement rather than coinciding with it is rarely possible to a standard that satisfies a sceptic. Track recurrence of specific eliminated defects against their own prior history rather than claiming credit for movement in a plant-wide figure.

Some assets are at the end of an economic life, and replacement is the right answer. An asset that is obsolete, unsupported, undersized for its duty or worn beyond economic repair will keep consuming effort and keep failing. Recognising this is not defeat; continuing to fund improvement work on it is the waste. A programme should be willing to produce replacement recommendations as an output, and one whose only output is maintenance activity cannot reach that conclusion.

Reliability is not free, and it is not a universal good

Every increment of reliability costs something: effort, capital, redundancy, downtime for intervention, or specification premium at purchase. Above a certain point the cost of the next increment exceeds the consequence it avoids, and pushing past that point is misallocation rather than diligence. Reliability is a trade against cost and criticality, which is why the criticality ranking tells you where the trade sits. There is no defensible universal target, and any figure presented as a benchmark for the reliability you should achieve was made up for a different organisation.

The idea to walk away with

Equipment reliability is a property with stated conditions and a stated period, determined first by design and selection, then by installation and commissioning quality, then by how the asset is operated, and only then by maintenance. Programmes that stay at the maintenance end work the weakest of the four levers, which is why so many run for years with high compliance and unchanged failure rates. The sequence that works is unglamorous. Make failures visible through consistent coding and honest downtime capture. Rank assets coarsely by criticality so you can legitimately do less on most of them. Get precision maintenance right, because you are already paying for the work. Eliminate recurring defects and verify they stopped. Review PM task content against real failure modes. Add condition monitoring later rather than first. Feed what you learn back into specifications. And name an owner with protected time.

Final thoughts

The most useful question to ask about a reliability programme is not which methodology it uses or what it has bought. It is: what has this programme changed outside the maintenance department. If the answer is a specification clause, a commissioning acceptance criterion, an operating procedure or a purchasing requirement, it is working the levers that set reliability. If the answer is more tasks, tighter compliance and a new dashboard, it is managing the consequences of reliability decisions taken elsewhere rather than improving them.

Two closing cautions. Be suspicious of any reliability target quoted without conditions and a period, including internal ones, because a claim that cannot be falsified cannot be managed. And resist the instinct to buy your way to the answer: the levers with the best return cost attention and discipline rather than capital, which is not a comfortable message for a business case but is the one that holds up.

For the wider discipline this article sits under, start from the reliability engineering complete guide. Primary sources worth going to directly: the IEC for the dependability vocabulary and the RCA standard, ISO for the reliability and maintenance data, condition monitoring and asset management families, and SAE International for the RCM evaluation criteria. All are voluntary, most paywalled, none law anywhere by itself; they carry weight through contracts, tenders and acceptance tests, which in practice is weight enough.

Disclosure

Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.

Building a reliability improvement programme?

Independent advisory on sequencing a reliability programme against your actual data maturity, failure data foundations, criticality-led effort allocation and the governance that keeps it from reverting to firefighting. 22+ years across utilities, oil and gas, manufacturing, government and facility operations.

Book a conversation

Related reading: What is reliability engineering, Reliability vs availability vs maintainability, Reliability metrics: MTBF, MTTR and availability, Equipment criticality analysis, Failure modes and how to analyse them, Root cause analysis methods, The bathtub curve explained, Failure codes: problem, cause, action, RCM: a practical introduction.

Muhammad Abbas

CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.

Work with me
MAbbaz.com
© MAbbaz.com