"Availability is down, so reliability is down." That sentence is a fixture of operations reviews, and it is wrong often enough to be dangerous. Availability can fall while reliability is untouched, because the asset is breaking no more often than before, it is simply taking four times as long to get back. It can also hold perfectly steady while reliability quietly collapses, because a fast crew and a full spares shelf are absorbing a rising failure count nobody has noticed. The three words describe three different properties of a system, and confusing them is why so many improvement programmes push on the wrong lever for a year and then wonder why the number has not moved.
The message up front: reliability is how often it breaks. Maintainability is how quickly and easily it can be put back, and it is mostly a property of the design rather than of the crew's effort. Availability is the outcome of those two, plus logistics and organisation. You cannot manage an outcome directly, which is why "improve availability" is a useless instruction on its own. The useful instruction names which input you are attacking, and that depends on which one is dominating your losses.
1. Why the three terms get tangled
Part of the tangle is that the industry has a formal vocabulary for all three and largely ignores it. The terms sit in the dependability vocabulary published as IEC 60050-192:2015, "International electrotechnical vocabulary, Part 192: Dependability", which supersedes the 1990 edition numbered IEC 60050-191. So the common claim that these words have no agreed definitions is false. In Europe, EN 13306:2017 "Maintenance, maintenance terminology", published by CEN with no ISO twin, defines the surrounding maintenance vocabulary.
The honest observation is that a published vocabulary has not settled the argument in practice. Practitioners still disagree about where the boundaries fall: whether waiting for a spare part belongs to maintainability or to logistics, whether a planned shutdown is unavailability, whether a standby asset is available. Those disagreements are not resolved by pointing at a clause, partly because the documents are paywalled and almost nobody in the room has read them. What resolves them is agreeing locally, in writing, using the standard terms as the anchor rather than inventing house language.
The second source of tangle is that only one of the three is routinely reported. Availability appears on the monthly pack; reliability and maintainability usually do not, or they appear disguised as MTBF and MTTR without anybody naming what those figures are properties of. The one number on the report is the composite, and the two that could explain it are missing.
2. Reliability: the probability of running without failure
Reliability is the probability that an item performs its required function, without failure, for a stated period, under stated conditions. Every clause in that sentence does work.
- It is a probability, not a duration. Reliability is not a number of hours. It is a chance expressed against a period: the probability that this pump runs the next 720 operating hours without a functional failure. MTBF is a related summary figure, but an average, not the probability itself.
- "Required function" has to be defined. A chiller that runs but delivers only 60 percent of its rated duty has failed functionally even though it is turning. Without a written function and performance standard you cannot say whether a failure occurred, and so cannot measure reliability at all.
- "Stated period" makes it comparable. Reliability over a week and over a year are different numbers for the same asset. Quoting it without the interval is meaningless.
- "Stated conditions" is the clause most often dropped. The same pump has different reliability handling clean water at steady duty than abrasive slurry with frequent starts. Reliability is conditional on duty, environment and how the asset is operated, not a badge on a nameplate.
Reliability is therefore a property of design, manufacture, installation, operating regime and the maintenance actually performed. That wide ownership is why reliability improvement is rarely a maintenance-only project. A pump that fails every six weeks because it is cavitating is not a maintenance problem, it is a hydraulic design or operating problem that maintenance is paying for, and crews are regularly performance-managed for a failure rate driven entirely by an upstream process change nobody told them about.
Reliability interventions are all about failure frequency: eliminate the failure mode, change the design or the duty, improve installation precision and alignment, fix lubrication and contamination control, apply condition monitoring, and use structured failure-mode analysis to find out what is actually breaking. For the metric mechanics behind all of this, the reliability metrics pillar on MTBF, MTTR and availability is the companion piece to this article, and the deeper treatments sit in MTBF explained and the failure rate formula and calculation.
3. Maintainability: how easily it can be restored
Maintainability is the ability of an item to be retained in, or restored to, a condition where it can perform its required function, given stated conditions of use and stated maintenance resources and procedures. In plain terms: how quickly and easily can this thing be put back once it has broken.
This is the concept most badly mangled in day-to-day use, and the mangling always takes the same form: maintainability is treated as a measure of the maintenance team's effort. It is not. It is overwhelmingly a property of the asset and its installed context. A crew cannot be diligent enough to overcome a motor that has to be craned out through a hole 50 millimetres too narrow, or motivated enough to overcome a controller with no diagnostic readout, or well-trained enough to overcome a machine whose bearing housing requires the whole drive train to be stripped to reach it. What actually determines maintainability:
- Physical access. Can a technician reach the maintainable item without scaffolding, confined-space entry or removal of unrelated equipment? Access is designed in at layout stage, and it is among the largest determinants of repair duration in practice.
- Modularity. If the failed element swaps out in twenty minutes and gets repaired off-line, downtime is short. If the repair must happen in situ, it is long.
- Diagnostics. Fault indication, error codes, test points, instrumentation that says which component failed rather than that something did. Diagnosis is often the largest single block of downtime and the least visible in reporting.
- Standardisation and parts commonality. Fewer distinct models, fewer distinct spares, common fasteners and interfaces. This shortens repairs, reduces stock value and widens technician competence at once.
- Documentation and tooling. Accurate drawings, parts lists that match what is installed, a written procedure, and no dependence on calibration rigs or software only the vendor holds. Documentation that disagrees with reality adds hours.
- Skill breadth required. If the repair needs a specialist who visits monthly, the asset is not very maintainable however simple the physical task is.
Notice that six of those seven are fixed before the asset is switched on for the first time. For the metric that tries to summarise maintainability, and its very real ambiguity, see MTTR meaning, formula and how to improve it.
The test that separates reliability from maintainability
Ask two questions about a downtime event. First: could a different design, duty or maintenance regime have stopped this from breaking at all? That is a reliability question. Second: given that it broke, could a different design have got it back faster? That is a maintainability question. The same outage almost always has an answer to both, and the two answers point at different budgets and different departments. A programme that only ever asks the first question leaves half the available improvement on the table.
4. Availability is a result, not a lever
Availability is the ability of an item to be in a state to perform as required: the proportion of time the asset is actually able to do its job. Crucially, it is not an independent property at all. It is what emerges when reliability, maintainability, logistics and organisation interact, and the causal shape is worth drawing explicitly:
↓
MAINTAINABILITY (how long each stop lasts, by design)
↓
LOGISTICS + ORGANISATION (parts, labour, permits, shifts)
↓
AVAILABILITY (the outcome you report)
Everything on the top three lines is something you can specify, buy, design, staff or change. The bottom line is arithmetic. This is why "improve availability by two points" is not an instruction, it is a wish: it contains no information about what anybody should do on Monday morning. The instructions that are actionable all name an input: reduce the failure frequency of this failure mode, cut the diagnosis time on this asset class, stock this part locally, extend crew cover into the night shift.
The same logic explains a recurring pattern in performance regimes. When availability is the only contracted or bonused number, the party measured pushes on whichever input is cheapest for them to move, which is usually not the one that matters most. If parts are the constraint and the contractor cannot influence stores, pressure on availability becomes pressure on technicians, producing rushed repairs, repeat failures and a decline in the reliability nobody was watching.
5. The three concepts side by side
This is the table I would put on the first slide of any reliability induction. The last column is the one that settles most arguments, because most misuse comes from asking a concept a question it cannot answer.
| Concept | What it describes | What it is a property of | How you influence it | What it does NOT tell you |
|---|---|---|---|---|
| Reliability | Probability of performing the required function without failure over a stated period, under stated conditions | Design, manufacture, installation quality, operating duty and environment, plus the maintenance actually done | Eliminate dominant failure modes, correct duty and environment, precision installation, lubrication and contamination control, condition monitoring, design change or redundancy | How long an outage will last, or how much of the year the asset will be usable |
| Maintainability | How quickly and easily the item can be retained in, or restored to, working condition | The asset and its installed context: access, modularity, diagnostics, standardisation, documentation, tooling | Mostly specification and selection: design for access and modularity, demand diagnostics, standardise models and spares, require usable documentation. Retrofit is possible but expensive | How often the asset will need restoring, or whether the parts and people will be there when it does |
| Availability | The proportion of time the item is in a state to perform as required | Nothing on its own: it is the compound result of reliability, maintainability, logistics and organisation | Only indirectly, by changing one of its inputs. There is no direct availability lever | Which of its inputs is causing the loss, or what the consequence of the lost time was |
The formal home for this family of terms is IEC , which publishes the dependability vocabulary, with the European maintenance terminology coming from CEN-CENELEC . If you are writing definitions into a contract or a corporate standard, anchor them to the published terms, because house definitions do not survive a change of contractor.
6. Two assets, identical availability, opposite problems
This is the most useful idea in the article, and the clearest demonstration of why availability on its own is a poor management instrument. Consider two assets over a year. The numbers are my own illustration, chosen to show the shape of the argument, not a benchmark or a target.
Asset A fails roughly twice a year, and each failure is a major event: a specialist flown in, a component ordered, the asset out for about five days. Asset B fails roughly weekly, and every failure is trivial: a technician clears a jam or resets a drive and it is running again in twenty minutes. Suppose the annual downtime totals land close to each other.
On a report showing only availability these two look nearly identical. In the plant they are nothing alike. Asset A is a continuity problem: when it goes, it goes for a week, and that is a business-interruption exposure spares strategy and redundancy design have to answer. Asset B is a chronic nuisance: it destroys crew productivity through constant interruption, trains operators to accept abnormal behaviour as normal, fills the backlog with low-value work orders, and is almost certainly a defect nobody has eliminated because each instance is too small to escalate.
They need opposite interventions:
| Asset A: rare but long | Asset B: frequent but short | |
|---|---|---|
| Reliability | Relatively good. Failures are infrequent | Poor. It fails constantly |
| Maintainability | Poor. Long diagnosis, difficult access, specialist skills, long-lead parts | Good. Fast to diagnose and reset by the on-site crew |
| Availability | Can be similar to B | Can be similar to A |
| What the operation actually feels | Occasional severe business interruption; long recovery; reputational and contractual exposure | Constant disruption; crew time fragmented; operators normalise faults; backlog fills with trivia |
| Correct primary intervention | Attack downtime duration: stock the critical spare locally, pre-stage the repair procedure, build in redundancy or a bypass, improve diagnostics, negotiate a response commitment, consider a module-swap strategy | Attack failure frequency: root-cause the recurring failure mode and eliminate it, correct the operating conditions or the design defect causing it, stop treating repeat resets as a fix |
| The wrong intervention | Pressuring the crew to work faster on a repair whose duration is set by parts lead time and access | Buying more spares and more crew so the resets get even faster, which makes the defect permanently affordable and therefore permanent |
| Where the budget sits | Spares capital, redundancy or capital modification, contract terms | Engineering investigation and a design or process fix |
The wrong-intervention row is worth dwelling on. Improving maintainability on Asset B makes the symptom cheaper without touching the cause, and cheap symptoms stop getting escalated. It is common for sites to become genuinely proud of how fast they can reset a machine that should never have needed resetting. Chronic short-duration failures also hide inside the work-order backlog rather than the downtime report, which is one reason backlog and downtime tracking should be read alongside availability, not after it.
7. Maintainability is designed in, and retrofitting it is expensive
Go back to the list in section three. Each item is settled at design, specification, selection or installation, and each becomes dramatically more expensive to change afterwards. Once the plant room is built, the pump is on its plinth and the pipework is welded, creating the access that should have been designed in is a capital project, not a maintenance improvement.
So maintainability belongs in procurement and design review, not in the operations improvement plan. That is an uncomfortable conclusion for a maintenance manager, because the decisions governing repair durations for the next twenty years are usually made by a project team that has moved on before the first failure. The lever, where it exists, is getting maintenance into the specification and the design review.
What I would put into a specification, in order of how much difference it makes:
- Maintenance access as an explicit design requirement. Stated clearances around every maintainable item, defined removal routes for the heaviest replaceable component, lifting provision where needed. Review it against the layout drawings before construction, not after.
- Modularity and diagnostics. Require the vendor to state which items are field-replaceable and how long each swap takes with named tools, plus fault indication that identifies the failed element and data output readable without proprietary tooling you do not own.
- Standardisation and spares commonality. Constrain the models, drives, controllers and fasteners in use, and insist on parts shared across the asset population rather than model-unique items. The cheapest maintainability intervention available, and usually surrendered to per-project purchasing decisions.
- Documentation as a deliverable with acceptance criteria. As-built drawings, parts lists matching the installed configuration, maintenance procedures and a spares list with lead times. Tie payment to it or you will not get it.
- No proprietary lock on routine maintenance. If a routine task needs vendor software, a vendor key or a vendor visit, your maintainability is a commercial variable outside your control.
Where you have inherited poor maintainability and cannot change the design, the realistic responses are compensatory rather than curative: hold the critical spare locally, pre-write and rehearse the repair procedure, cross-train enough people, and where the consequence justifies it add redundancy. Those are genuine improvements, and they are also more expensive, forever, than the design decision that would have avoided the problem.
The limitation of this whole framing
Separating the three concepts is clarifying, but the boundaries genuinely blur at the edges. Is a repair slowed by a missing spare a maintainability problem or a logistics one? Is an asset that fails because a poor design forced an intrusive repair a reliability or a maintainability failure? Is redundancy a reliability measure or an availability measure? There are defensible answers, and they differ between texts and industries. The value of the distinction is diagnosis, working out which lever to pull, not adjudicating every case. If you find yourself in a long argument about which category an event belongs to, the categorisation has stopped earning its keep. Record what happened in enough detail that it can be counted either way, and move on.
8. The contributors that sit outside both
Here is the part most conceptual treatments understate. A substantial share of real-world unavailability is caused by neither the asset's reliability nor its maintainability, but by the organisation around the asset. These are usually grouped as logistic and administrative delays, and on many sites they are the largest single block of downtime.
- Parts availability. One of the most common constraints. The repair takes two hours and the part takes six weeks. No design change and no crew training affects that number. It is a stocking policy, a criticality decision and a supplier arrangement, which is where spare parts and MRO inventory strategy becomes a reliability discipline rather than a stores one.
- Labour availability and shift coverage. Whether a competent person is on site when the failure happens. A fast repair that begins nine hours later because the failure happened at 6pm on a Thursday is a fast repair and a long outage. Weekends, public holidays, shutdown seasons and production windows all bite here; in the Gulf, summer peak-load restrictions on taking certain plant offline are a real constraint on repair timing.
- Permit and access delays. Isolation, permit to work, confined-space entry, hot-work authorisation, and simply getting operations to release the asset. On higher-hazard plant these are properly non-negotiable, but they are frequently slower than the safety requirement itself demands, because the administrative process has never been examined.
- Specialist mobilisation. If the fix needs a vendor engineer, your availability is partly governed by their travel schedule and contract terms.
- Detection delay. The time between the asset losing function and anybody noticing. On unmonitored and remote assets this can dwarf the repair itself, and it is neither a reliability nor a maintainability property. It is an instrumentation decision.
These contributors are the usual explanation for the gap between what an asset is capable of and what it delivers. When you decompose downtime, keep them separate in the data, because lumping a six-week parts delay into "repair time" makes a procurement problem look like a maintenance problem and guarantees it gets attacked in the wrong place. Most mainstream maintenance systems can capture this through wait-state or delay coding and very few organisations use it; if you are setting up a system, the CMMS buyer's introduction covers the data-capture foundations this depends on.
9. Inherent versus operational availability, conceptually
Everything in the previous section explains why availability comes in more than one flavour, and why the flavour matters commercially.
Inherent availability is a design property. It considers only reliability and corrective maintainability, assuming parts are instantly available, a skilled technician is standing there, and no permit, access or administrative delay exists. It is a laboratory figure, legitimately useful for what it is designed for: comparing two candidate designs on equal terms with the site-specific noise stripped out. Operational availability includes everything: every delay, every wait for a part, every hour the asset sat idle because the permit had not been signed. It is what the operations manager experiences and what the production figures reflect.
The commercial consequence is worth saying plainly: vendors quote the first and operators live the second. That is not necessarily dishonest. The inherent figure is the only availability a manufacturer can legitimately claim, because the delays separating it from the operational figure are properties of your organisation, not their equipment. The error is on the buying side, where a datasheet figure gets written into an operational target or a contract KPI without adjustment and the gap then gets blamed on the maintenance team. The same equipment in a site with good stores, round-the-clock cover and a clean permit process will deliver close to its inherent figure. In a site with a four-week parts lead time it will not, and no pressure on the crew changes that.
There are further formal variants, notably an achieved availability sitting between the two by including preventive work but still excluding logistic delay. The arithmetic and the ways the denominator gets quietly narrowed are covered in the reliability metrics pillar, and I am deliberately not repeating the calculations here. The conceptual point is enough: there is more than one availability, they differ by what delays they admit, and the difference between them measures your organisation rather than your equipment.
10. RAM analysis: what it is and when it is done
Once you accept that availability is the compound result of reliability and maintainability, the natural engineering question is whether you can predict it before the plant exists. That is what RAM analysis does. RAM stands for reliability, availability and maintainability, and a RAM study estimates the availability a proposed system will deliver, given assumptions about how often each element fails and how long each takes to restore.
In outline it breaks the system into functional blocks and maps how they combine, which elements are in series so any one failing stops the function and which are redundant; assigns each block a failure and a restoration characteristic from generic industry data, vendor data or the operator's own history; layers on the logistic assumptions, spares holding, crew availability and mobilisation times, because these drive the result as strongly as the equipment data; computes the resulting system availability over the design life; and ranks the contributors. That ranking, not the headline number, is the genuinely valuable output.
When it is done matters. RAM analysis belongs to design and capital projects: front-end engineering, concept selection, comparing plant configurations, sizing redundancy and initial spares holding, and supporting a contractual availability guarantee. It is not a routine operations activity. Commissioning one on an existing plant you cannot modify produces a document rather than an improvement, because the decisions it informs have already been made.
Its value is comparative rather than absolute. A model saying a design will achieve some specific availability figure is stating the consequence of its input assumptions, which carry wide uncertainty. A model saying configuration one will outperform configuration two, and that the dominant contributor in both is the single non-redundant element in the middle, is telling you something robust and actionable.
If you want the failure history your own organisation collects to be usable in this kind of analysis, the taxonomy and data-format reference is ISO 14224:2016, "Petroleum, petrochemical and natural gas industries, collection and exchange of reliability and maintenance data for equipment", published by ISO . It is written for oil and gas but its equipment taxonomy is widely borrowed elsewhere. One frequently muddled clarification: OREDA is a proprietary members-only reliability database, not a standard. ISO 14224 is the standard that grew out of that work.
Where RAM analysis disappoints
A RAM model is only as good as its failure data and its independence assumptions, and both are usually weaker than the output's precision suggests. Generic industry failure data may not describe your duty, environment or maintenance quality. Redundancy calculations assume elements fail independently, which duty and standby pairs sharing a power supply, a control system, an environment and a batch of spares rarely do. And a model built during a project is seldom revisited once real operating history exists, which is exactly when it would become accurate. Treat the original as a design comparison, not a forecast of what you will experience.
11. How to use the distinction in practice
The distinction earns its keep in one repeatable way: it turns "availability is bad" into a diagnosis. The sequence I would follow:
- Decompose the loss into frequency and duration. Count failure events and total downtime hours separately. Frequency is a reliability signal, duration a maintainability and logistics signal. Then split duration into detection, mobilisation, diagnosis, waiting for parts or permits, active repair and handback. The largest block is where the intervention belongs, and it is very often not the active repair.
- Name the lever, not the target. Never write an objective that says only "raise availability". Write which input you are attacking and by what mechanism, so the action is testable, and report failure count, downtime decomposition and availability on the same page.
- Route the finding to the right owner. Frequency problems usually go to engineering and operations as well as maintenance. Duration problems split between design and specification, stores, and the permit process. Very few belong solely to the maintenance crew, and sending them all there is the commonest management error in this area.
- Set targets against your own baseline. There is no credible universal benchmark for availability, reliability or repair duration. The figures depend on asset mix, duty, redundancy design, environment and the boundary definitions you chose. Trend your own numbers under stable definitions and treat any external comparison as a conversation-starter.
Practical depth on the reliability side sits in equipment reliability and how to improve it, and the wider discipline above all of this is covered in the complete guide to reliability engineering.
The idea to walk away with
Reliability is a property of how often the asset breaks. Maintainability is a property of how easily it can be put back, and it lives in the design rather than in the crew's commitment. Availability is neither: it is the arithmetic consequence of both, plus the parts, people, permits and shift patterns around them. Because it is a consequence it cannot be managed directly, and knowing which input dominates your particular loss is the entire diagnostic value of keeping the three concepts apart.
The corollary is the two-assets case. Identical availability can describe an asset that fails rarely and stays down for days and one that fails weekly and is back in twenty minutes, and those two need opposite responses. The first needs shorter downtime, through spares, redundancy, diagnostics and rehearsed procedures. The second needs fewer failures, through root-cause elimination of a defect that fast recovery has been quietly making affordable. Availability alone cannot distinguish them. Reliability and maintainability, reported separately, cannot fail to.
Final thoughts
The three words have a formal home in IEC 60050-192:2015, with the surrounding maintenance vocabulary in EN 13306:2017, and knowing that is worth something even though it will not stop the arguments. What stops the arguments is a local agreement, anchored to the published terms, about what counts as a failure, which delay categories exist, and which of them get reported against whom.
If there is one habit to change after reading this, it is to stop accepting availability as a standalone number. When the monthly pack shows availability down two points, the only useful next question is whether the asset is failing more often or staying down longer, and if longer, which part of the down period grew. Ask that reflexively and the conversation moves from blame to diagnosis. Most of the real gains here come not from a new technique but from splitting one number into its parts and finding the problem had been in stores, or in the permit process, or in a design decision taken fifteen years earlier, the whole time.
Disclosure
Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.
Availability not moving despite the effort?
Independent advisory on decomposing availability into reliability, maintainability and logistics, downtime data capture, maintainability requirements for equipment specification, and the KPI set that makes the diagnosis visible. 22+ years across utilities, oil and gas, manufacturing, government and facility operations. No equipment vendor arrangements.
Book a conversationRelated reading: What is reliability engineering: a complete guide, Reliability metrics: MTBF, MTTR and availability, MTBF explained, MTTR: meaning, formula and how to improve it, Failure rate: formula and calculation, Equipment reliability and how to improve it, Maintenance backlog and downtime tracking, Spare parts and MRO inventory in a CMMS.
Muhammad Abbas
CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.
Work with me