mail@mabbaz.com Abu Dhabi, UAE

Reliability Metrics · Maintenance Performance · CMMS

MTTR: Meaning, Formula and How to Improve It

MTTR is one of the simplest calculations in maintenance and one of the most argued about. The arithmetic takes a line. The hard part is agreeing what "repair" includes, when the clock starts, and when it stops. This guide covers the meaning, the formula, the definitional problem that decides the number, how to decompose MTTR into stages you can actually improve, and how the metric gets gamed.

Muhammad Abbas September 27, 2026 ~18 min read

Ask five people in the same maintenance organisation what MTTR was last month and you will often get five answers, all of them calculated correctly. That is not a data quality failure, or at least not only that. It is a definitional failure. MTTR is an average of a duration, and the duration has two boundaries that nobody agreed on: the moment the clock starts and the moment it stops. Shift either boundary and the same set of real events produces a different number. Most MTTR disputes I have been called into were never disputes about arithmetic at all.

The message up front: MTTR is total repair time divided by the number of repairs. That part is trivial. The number only means something once you have written down, in one sentence, which event starts the clock, which event stops it, and which waiting periods are inside the span. Until then you have a figure, not a metric. And because different organisations answer those three questions differently, comparing your MTTR to somebody else's is close to meaningless.

1. What MTTR means

MTTR stands for mean time to repair. It is the average amount of time it takes to restore a failed asset to working order, measured across a set of repair events over a defined period. It is a maintainability measure in its strictest reading: it describes how quickly and easily a failed thing can be put right, which is a property of the asset, the design, the spares position, the documentation and the crew, all together.

That is the textbook sentence, and it hides the whole problem, because "restore a failed asset to working order" is not a single observable event. A pump fails at some point in the night. An operator notices at shift handover. A work request is raised twenty minutes later. A technician is assigned, travels, diagnoses, discovers the seal kit is not in stores, waits two days, fits it, runs the pump up, and operations formally accepts it back the following morning. Which of those moments is the start of "time to repair"? Which is the end? Every one of those choices is defensible, and they produce durations that differ by orders of magnitude from one identical event.

There is a formal vocabulary for this family of terms. Dependability terminology, including the mean-time metrics, lives in IEC 60050-192:2015, "International electrotechnical vocabulary, Part 192: Dependability", which supersedes the 1990 edition IEC 60050-191 that older documents still cite. IEC publishes the vocabulary and makes its terminology browsable through Electropedia . The honest point to make about it is this: a formal vocabulary exists, and industry practice still disagrees about the clock anyway. Contracts, CMMS reports, vendor datasheets and management packs all use "MTTR" for different spans. That gap between the vocabulary and the practice is the source of most MTTR arguments, and no amount of citing a standard in a meeting resolves it. Writing your own definition down does.

If you want MTTR placed alongside its siblings rather than examined on its own, the comparative treatment is in the reliability metrics pillar on MTBF, MTTR and availability, which sets out how the metric set fits together and how availability is computed from it. This article stays on MTTR itself and, in particular, on how to move it.

2. The arithmetic, with an illustrative example

The formula is a plain average:

MTTR = total repair time / number of repairs

Nothing more than that. It is a sum of durations divided by a count of events. An illustrative example, with numbers I have made up purely to show the mechanics: suppose one chilled water pump suffered four failures during a quarter, and the repair durations, measured on a definition you have already agreed, were 3 hours, 1 hour, 14 hours and 2 hours. The total is 20 hours across 4 repairs, so MTTR is 5 hours.

Look at what that single figure has done. Three of the four repairs were short. One was long, almost certainly for a reason that had nothing to do with repair skill, and it has pulled the average above every value except itself. A mean over a small number of events with a skewed distribution is a weak summary, which is why I would always ask to see the individual durations next to the mean, and often the median as well. For the full calculation walkthrough, including period selection, how to treat repairs that span period boundaries, and how to handle multiple assets in one figure, see the MTTR formula and calculation guide. This article deliberately does not repeat it.

What the mean hides

Repair durations are rarely normally distributed. They cluster low with a long tail, and the tail is usually caused by logistics rather than by the repair itself. An MTTR that improves because one bad event fell out of the reporting window has not improved anything. Report the count, the mean, the median and the longest event together, or the number will mislead you and your management.

3. The problem that decides the number: what "repair" includes

This is the spine of the whole subject, so it is worth being slow about. An MTTR definition has to answer three questions, and each one has several defensible answers.

Question one: when does the clock start? The candidates, in the order they occur:

  • At failure. The moment the asset stopped performing its function. This is the operationally honest start, and for most assets it is unknown. Unless the asset is instrumented and the trip is timestamped by a control system, nobody can say when it actually failed, only when somebody noticed.
  • At detection or notification. The moment a human or a system raised the alarm. Usually the first timestamp that genuinely exists. It quietly excludes detection delay, which on unmonitored assets can be the largest single component of the outage.
  • At work order creation or assignment. Convenient, because every system records it, and it excludes whatever administrative delay sat between the request and the work order. That delay is real downtime and it is now invisible.
  • At technician arrival or work start. This is the narrowest start. It gives you a clean measure of how long the repair took once someone was actually working on it, and it excludes everything the maintenance crew would argue is not their fault.

Question two: when does the clock stop? Fewer candidates, but they matter just as much:

  • At tools down. The technician has finished the physical work. Nothing has been proven yet.
  • At restart. The asset has been run up and is turning. On plant with a commissioning or stabilisation period, this is earlier than useful service.
  • At handback and acceptance. Operations has tested the asset and formally taken it back. This is the operationally true end and it is the one most likely to be missing from your data, because handback is a conversation rather than a transaction in most organisations.
  • At work order closure. Administratively tidy, operationally wrong. Work orders get closed days or weeks after the asset was back in service, and closure backlog will contaminate the metric badly.

Question three: which waiting periods are inside the span? Waiting for a spare part. Waiting for site or equipment access. Waiting for a production window or a shutdown slot. Waiting for a specialist contractor to mobilise. Waiting for a permit to be issued. Waiting for the equipment to cool, drain or de-energise before it is safe to touch.

These are the components that decide whether your MTTR is one hour or three days, and they are almost entirely outside the control of the person holding the spanner. A narrow definition that excludes them measures crew efficiency and asset maintainability, which is what the term was originally for. A broad definition that includes them measures what operations experienced, which is what the business cares about. Both are legitimate. Publishing one while your counterparty assumes the other is how contract disputes start.

The test to apply

Write your MTTR definition as one sentence naming two timestamp fields and an explicit list of what is included and excluded. If two analysts cannot independently produce the same number from the same work order history using that sentence, the definition is not finished. Do this before you agree a target, not after somebody disputes a report.

One clause on permits, because they come up constantly as a source of delay and the answer is jurisdictional. Permit-to-work practice is set by local law, by the site operator and by sector convention rather than by any international standard; the widely used reference is UK HSE guidance HSG250, "Guidance on permit-to-work systems", which is guidance rather than law and was written for the petroleum, chemical and allied industries, so it does not automatically describe how a hospital or a campus should run permits. It is freely published by the UK Health and Safety Executive . For how permit time interacts with the work order clock in practice, see the permit to work guide.

4. The clock-definition choices, side by side

The table below sets out the four common clock definitions, what each one includes, what it excludes, and what it is honestly good for. None of these is wrong. Choosing without knowing which one you chose is wrong.

Clock definition Start → stop Includes Excludes Fit for
Active repair only Work start → tools down Hands-on repair, on-site diagnosis once working Detection, admin, travel, all waiting, testing, handback Maintainability, crew productivity, task time standards
Response inclusive Notification → tools down Triage, assignment, travel, repair Detection delay, parts and permit waiting if paused, testing In-house crew performance where parts are held on site
Restoration Notification → handback accepted Everything after notification, including all waiting and testing Detection delay only Service contracts, SLA rectification windows, operations reporting
Full downtime Failure → handback accepted Detection delay plus everything above Nothing within the outage Availability calculation, business impact, loss accounting

My standing recommendation is to capture the timestamps that let you produce the narrow figure and the full-downtime figure from the same data, then report both with different names, never both as "MTTR". Use MTTR for the repair span you defined, and use mean down time for the full outage. That single naming discipline removes most of the confusion at no cost.

5. The metrics people conflate with MTTR

A large part of the ambiguity is that several genuinely different metrics share the same initials or nearly do, and they get used interchangeably in the same report.

  • Mean time to acknowledge: from notification to somebody confirming they have taken ownership. It measures the responsiveness of the intake and triage process, and nothing else. Useful, frequently the fastest thing to improve, and completely silent on how long the fix takes.
  • Mean time to respond: from notification to a technician being on site and starting work. This is the number most facilities and service contracts actually specify, often labelled MTTR in the contract. It is a mobilisation measure.
  • Mean time to recover or restore: from failure or notification to the asset being back in service. Broader than repair, and the one operations experiences.
  • Mean time to resolve: common in IT service management, extending to formal incident closure, which can be well after service was restored.
  • Mean down time: the average total unavailability per failure event. The cleanest name for the full-outage figure, and the one that belongs in an availability calculation.

Separating these is not pedantry, it is what makes the metric actionable. A single blended MTTR tells you the number is bad. Acknowledge, respond, repair and restore measured separately tell you where it is bad, and each of those four has a different owner and a different fix. If acknowledge time is poor you have an intake and rota problem. If respond time is poor you have a dispatch, coverage or travel problem. If active repair is poor you have a skills, tooling, documentation or asset-design problem. If restore time is poor while active repair is good, you have a logistics problem, almost always spares. One metric cannot distinguish those, and four can.

For the wider metric family, including how MTTR pairs with failure-frequency measures, see MTBF explained and MTTF, the formula and examples. MTTR also sits inside the availability component of OEE, which is worth understanding if your MTTR is being used to explain a production loss figure.

6. Decomposing MTTR by stage, so you can see where the time goes

A single MTTR figure cannot be improved because it does not tell you what to change. A staged decomposition can. The stages below are the ones I would set up in any maintenance system, each with the lever that typically moves it. Populate the timestamps once and every row becomes measurable.

Stage What it measures Timestamps needed Typical lever
Detect Failure to somebody knowing Failure time (instrumented) → detection Condition monitoring, alarms, operator rounds, meter thresholds
Notify and triage Knowing to the job being owned Detection → request raised → assigned Single intake channel, mobile reporting, clear escalation and rota
Mobilise Assignment to on-site work start Assigned → arrival → work start Zoned crews, shift coverage, route planning, van stock
Diagnose Finding the actual fault Work start → fault confirmed Failure history, drawings and manuals to hand, fault trees, skills
Wait for parts Logistic delay on materials Part requested → part in hand Critical spares policy, kitting, reorder points, supplier lead times
Wait for access or permit Delay on safety and availability Access requested → access granted Pre-approved routine permits, isolation planning, agreed outage windows
Repair Hands-on work Repair start → tools down Task procedures, special tools, second pair of hands, design for access
Test Proving the fix Tools down → test complete Defined acceptance tests, instrumentation, run-up procedure
Hand back Formal return to operations Test complete → accepted Named acceptor, mobile sign-off, no waiting for the morning meeting

When a site tells me its repairs are fast but availability is poor, the answer is usually in the waiting rows, not the repair row. The crew is competent and the seal kit is four days away. That distinction is invisible in a blended figure and obvious in a staged one, and seeing it is what redirects the improvement effort to where it will actually pay. Tracking it over time also connects directly to your backlog and downtime tracking, because a growing waiting-for-parts component and a growing backlog usually have the same root cause.

7. How to improve each stage

Improvement is stage by stage. There is no single intervention that lowers MTTR, which is why "reduce MTTR by a quarter" as a standalone objective produces activity rather than results. Taking the stages in order:

  • Detect faster. On critical assets, instrumentation and alarming remove detection delay almost entirely. On the long tail of non-critical assets, the cheap version is operator rounds and making it trivially easy for anyone on site to report a fault. Detection delay is the component most often excluded from the metric and most often the largest, which is a bad combination.
  • Shorten notify and triage. One intake channel, not five. A request that arrives by phone call to a supervisor who is on annual leave is the classic failure here. Clear out-of-hours escalation, an unambiguous owner for every priority level, and the ability to raise and assign from a phone. This stage is usually the cheapest to fix and it is often measured in hours.
  • Cut mobilisation. Zoning technicians by area, sensible shift coverage on the hours when failures actually occur, van stock for the parts consumed most often, and scheduling that does not require a technician to cross the site twice. On multi-site portfolios, travel is frequently the single largest controllable component.
  • Speed up diagnosis. This is where good failure history pays back. A technician who can see the last four faults on this asset and what fixed them diagnoses faster than one starting from nothing. Drawings, manuals, points lists and P&IDs accessible on a phone at the asset, rather than in a filing cabinet, remove a large slice of diagnostic time. Consistent failure coding is what makes the history usable in the first place.
  • Attack parts waiting. Often the biggest single lever on long outages, and commonly owned outside maintenance. Decide deliberately which spares are held for which critical assets, kit the parts for known repeat repairs, set reorder points from real consumption rather than guesswork, and know your actual supplier lead times rather than the quoted ones. The trade-off is explicit: holding stock costs money and idle downtime costs money, and for critical assets the second usually wins. The mechanics of that are in spare parts and MRO inventory in a CMMS.
  • Reduce access and permit waiting. Pre-approved permits for routine, well-understood tasks. Isolation plans prepared in advance for assets you know will need them. Agreed outage windows with operations rather than a negotiation each time. Keep the safety rigour and remove the administrative queue; those are separable, and conflating them is how sites end up either slow or unsafe.
  • Shorten the repair itself. Written task procedures for repeat repairs, the special tools actually available rather than theoretically owned, planning for two people where one cannot do it safely, and, over the longer term, feeding maintainability back into procurement so the next asset is specified with access and serviceability in mind. Design decisions made at purchase set a floor under your repair time for the asset's whole life.
  • Make testing deliberate. A defined acceptance test per asset class, known in advance, avoids the pattern where a repaired asset sits waiting for somebody to decide how to prove it works.
  • Remove handback friction. A named acceptor per area and the ability to sign off from a phone. I have seen assets sit repaired and unavailable overnight because the only person who could accept them back was on the day shift. That is pure administrative downtime and it is embarrassing once measured.
Where this approach does not help

Driving MTTR down is the wrong objective when the real problem is failure frequency. An organisation that has become excellent at fast repairs of the same recurring fault has optimised the symptom. On assets with repeat failures, the return on eliminating the failure mode dwarfs the return on repairing it faster, and a good MTTR can actively mask that by making a badly performing asset look well managed. MTTR is only meaningful read next to how often the asset fails.

8. Why cross-organisation MTTR benchmarking is close to meaningless

This follows directly from section three. If MTTR depends on which clock definition you chose, and organisations choose differently, then two MTTR figures from two organisations are two different measurements wearing the same label. Add the further differences, the asset mix, the criticality profile, the failure modes counted, whether minor faults are logged at all, whether contractor work is included, and the comparison carries almost no information.

I do not publish target or benchmark MTTR figures for any industry or asset class, and I would treat any that you are shown with suspicion unless the source states its clock definition, its asset scope and its counting rules. Where a published figure exists without those three things, it is a number, not a benchmark.

The useful comparison is against yourself. Fix the definition, establish a baseline from your own history, and measure the trend in the staged components. A site that can show its parts-waiting component halving over two quarters has demonstrated something real. A site that can show its MTTR is lower than an industry figure it found online has demonstrated nothing. Targets should be set against your own baseline for exactly this reason.

9. How MTTR gets gamed

Most MTTR distortion is not fraud. It is the predictable behaviour of people responding to a metric that can be improved without improving anything. The patterns to watch for:

  • Narrowing the clock quietly. Moving the start from notification to work start, or the stop from handback to tools down, improves the figure overnight with no operational change. Check whether a sudden improvement coincided with a report being rebuilt.
  • Pausing the clock. Legitimate in principle, for genuine waiting on a third party, and easily abused. If work orders can be put on hold and held time is excluded, hold status becomes the place where downtime goes to hide. Audit the held time as a metric in its own right.
  • Splitting one repair into several work orders. A three-day outage recorded as four short work orders lowers the mean and raises the count. Look for multiple work orders on the same asset with adjacent timestamps.
  • Not recording the small stuff. If quick fixes are logged and long ones are handled informally as projects, the recorded population is biased low. Conversely, if only major failures are logged, it is biased high.
  • Closing early, finishing later. Marking the work complete at the point the technician leaves, then returning to finish, produces a clean metric and a still-broken asset.
  • Reclassifying failures as planned work. A corrective job recorded as a scheduled task leaves the failure population entirely.
  • Backdating timestamps at closure. Times entered from memory days later cluster on round numbers. If your arrival times are overwhelmingly on the hour, they are estimates.

The defence is not surveillance, it is measuring the components separately so that no single number can be improved in isolation, and reporting alongside MTTR the things that would have to move with it if the improvement were real: unplanned downtime hours, failure counts, emergency work percentage. If MTTR improves while downtime hours are flat, something definitional changed rather than something operational.

10. The data discipline that makes all of this possible

Everything above depends on timestamps, and this is where most organisations fall short. The fields exist. IBM Maximo, SAP PM, Hexagon EAM, Planon and the mid-market tools such as Fiix, Limble, eMaint, UpKeep and MaintainX all carry several of the relevant timestamps, some automatically, some requiring configuration. The problem is rarely the software. It is that two or three of the fields are populated reliably and the rest are populated when somebody remembers.

The minimum set worth insisting on: detection or reported time, request raised, assigned, arrived on site, work started, work completed, and handback accepted, plus start and end times for any hold, with a reason code on the hold. That last item is what turns "we waited" into "we waited four days for a seal kit", which is the difference between a complaint and an improvement case. Software matters here only because stage decomposition is impossible without automatically captured times; the general principles of what a maintenance system should capture are covered in the CMMS introduction and, for the work order layer specifically, in work order software for maintenance teams.

Two practical cautions. First, timestamps entered manually at closure are not measurements, they are recollections, and a metric built on them will be smooth, plausible and wrong. Capture times as events happen, from a mobile device at the asset, or accept that the numbers are indicative. Second, do not add nine mandatory fields at once. Add the two that unlock the component you most suspect is your problem, prove they get populated, then extend. A shorter list that is completed honestly beats a complete list that is filled in from memory.

The idea to walk away with

MTTR is arithmetic sitting on top of a definition, and the definition does all the work. Total repair time divided by number of repairs is the easy half. Naming the start event, the stop event and the included waiting periods is the half that decides whether the number means anything, and it is the half most organisations skip. Once you have written that sentence down, MTTR stops being a number people argue about and becomes a measurement people can act on.

The action itself comes from decomposition. Detect, notify, mobilise, diagnose, wait for parts, wait for access, repair, test, hand back. Each stage has a different owner and a different lever, and a blended figure hides all of them. Improvement follows from finding which stage holds your time and pulling the lever that matches it, not from setting a target on the aggregate and hoping.

Final thoughts

If I could change one habit in how MTTR is used, it would be to stop treating it as a score and start treating it as a diagnostic. A score invites comparison with other organisations, which the metric cannot support, and invites gaming, which it makes easy. A diagnostic invites the question that actually improves reliability: of the total time this asset was out of service, how much was detection, how much was waiting, and how much was repair?

And keep MTTR next to failure frequency. Fast repair of a fault that keeps recurring is competence applied to the wrong problem. Organisations that make real progress do two things in parallel: they define their clock properly so the number can be trusted, and they keep asking why the asset was in front of a technician again at all.

Disclosure

Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.

Need your MTTR definition to survive an audit?

Independent advisory on maintenance metric definitions, timestamp design in CMMS and EAM, SLA clock wording and the reporting that proves improvement is real. 22+ years across utilities, oil and gas, manufacturing, government and facility operations.

Book a conversation

Related reading: Reliability metrics: MTBF, MTTR and availability, MTTR formula: how to calculate MTTR, MTBF explained, MTTF: formula and examples, Maintenance backlog and downtime tracking, Spare parts and MRO inventory in a CMMS.

Muhammad Abbas

CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.

Work with me
MAbbaz.com
© MAbbaz.com