Ask a maintenance manager how big the backlog is and you will usually get a number of open work orders. Ask what that number means and the conversation stalls. Is a work order raised this morning backlog? Is one waiting eleven months for a part that was never ordered backlog? Is a statutory inspection scheduled for next quarter backlog? In most organisations I have worked with the answer to all three is yes, because the report is simply a count of everything not closed. That number is useless for planning and quietly corrosive, because when management asks why it keeps rising nobody can give an honest answer. Downtime is the same story with different symptoms. Both are fixable with definitions and habits rather than software.
The message up front: backlog is not everything open, it is the work that is ready to be scheduled, and it is measured in crew-weeks rather than job count. Zero backlog is a warning sign, not an achievement. Downtime is only useful when every hour carries a reason code and a clear owner, and when maintenance and operations agree in advance on what the clock is measuring. Get the definitions right and both numbers become management tools. Get them wrong and you have two reports that nobody trusts and everybody argues about.
1. What actually counts as backlog
The most valuable thing you can do to a backlog report is narrow it. Backlog, properly defined, is identified work that has been approved, scoped and resourced to the point where it could go on a schedule, and has not yet been executed. That deliberately excludes several categories of open work order most reports lump in.
- Not backlog: unapproved requests. A service request waiting for a supervisor to decide whether it is real work is a demand queue, not backlog. Measure it separately: a growing request queue is a planning problem, a growing backlog is a capacity problem, and conflating them hides both.
- Not backlog: future-dated planned work. A PM task generated for a due date three weeks out is forward schedule. Count it and your backlog rises and falls with the PM generation cycle rather than with anything meaningful about capacity. Overdue PM is a different measure and belongs with schedule compliance.
- Backlog with a qualifier: approved but blocked. A job waiting on a part, a permit, a shutdown window or a client approval is genuinely backlog, but it is not schedulable this week and should never sit in the same bucket as ready work.
- Backlog, unambiguously: approved, scoped, resourced, ready, not yet done. The number the planner works from, and the number that tells you whether crew size matches workload.
Work order types matter here, because a corrective job raised from a PM inspection behaves differently in the backlog from an emergency callout that never queues at all. If your type taxonomy is loose, your backlog definition will be too. The work order types in a CMMS guide covers the taxonomy that makes this segmentation possible.
The test I apply to any backlog report
Pick any line on the report and ask: could a planner put this job on next week's schedule and expect it to be done? If yes, it is ready backlog. If no, something is blocking it, and the report must say what. A backlog report that cannot answer that question for every line is a list, not a management tool.
2. Why crew-weeks beats job count
A count of 300 open jobs tells you nothing about capacity. Three hundred filter changes is a fortnight for two technicians. Three hundred jobs including four pump overhauls and a switchboard refurbishment could be six months. The unit that carries meaning is labour hours normalised against available capacity, expressed as backlog weeks, sometimes called crew-weeks. The calculation is deliberately simple:
÷ weekly available crew hours
Example structure (use your own figures):
8 technicians × 40 hours = 320 gross hours/week
less leave, training, travel, breaks, emergency response
= net wrench-time capacity available for planned work
↓
Backlog hours ÷ that net figure = backlog weeks
Two things about this calculation cause most of the arguments. First, the numerator depends on every backlog job carrying a labour estimate. If half your work orders have no estimated hours, the figure is fiction. A default estimate by job type is far better than a blank field, and it improves as actuals accumulate.
Second, the denominator must be net capacity, not headcount times 40. Deduct leave, training, toolbox talks, travel between buildings on a multi-site portfolio, and the share of capacity reactive work will consume whether you plan for it or not. Using gross hours is the most common way a genuinely unhealthy backlog gets reported as comfortable.
Report backlog weeks by trade, not just in total. A total of four weeks hides mechanical at two and electrical at nine. Total backlog is a board number; backlog by trade tells you where to hire, where to use a contractor, and where to stop accepting new work.
3. Healthy and unhealthy backlog, and why zero is bad
This is the part that surprises people outside maintenance. A backlog of zero is not an achievement. It means you have no buffer of prepared work, which in practice means your crew is oversized for the workload, your inspection regime is not finding defects, or your planners are pushing work straight to the floor without the preparation that makes it efficient. None of those is good news.
A healthy backlog is a queue of prepared, parts-available, scoped work that lets the planner build a full week of efficient schedules and lets a crew that finishes early pull the next job rather than stand idle. The conventional planning view, which matches what I see work in practice, is a ready backlog of a few crew-weeks per trade: enough to schedule from comfortably, not so much that the oldest work is stale. The bands I use when advising a team, as judgement rather than industry law:
| Ready backlog | What it usually means | What I would do |
|---|---|---|
| 0 to 1 week | No planning buffer. Crew works hand to mouth, idle between jobs, planners firefighting. Often a sign inspections are not generating defects. | Check that inspection routines raise corrective work. Review whether crew size exceeds workload. |
| 2 to 4 weeks | Generally healthy. Efficient weekly schedules, parts staged in advance, jobs kitted before release. | Hold it here. Protect the buffer from being raided for urgent work without a decision. |
| 4 to 8 weeks | Under pressure. Work is ageing, requesters start chasing, priority begins to be decided by whoever complains loudest. | Decide explicitly: add contractor capacity, or defer and cancel work through a proper review. Do not let it drift. |
| Over 8 weeks | Structural capacity gap, or a backlog full of work that should have been cancelled months ago. Usually both. | Run a full backlog review to strip out dead work, then size the residual gap honestly and take it to management with a cost. |
Treat those bands as a frame to argue with, not a benchmark. A single-site plant with predictable duty and a hard annual shutdown runs a very different profile from a multi-site facilities contract with SLA obligations. What matters more than hitting a band is that the number is calculated the same way every week, so the trend is real.
Where the backlog-weeks measure stops helping
Backlog weeks assumes labour is the binding constraint. On plenty of sites it is not. If your work is blocked mainly by parts lead time, by access to occupied space, or by a shutdown window that comes once a year, you can have two crew-weeks of backlog and still be structurally unable to clear it. In those environments backlog weeks is a secondary measure, and the segmentation in the next section is the primary one. Know which constraint actually governs your site before you make backlog weeks the headline.
4. Backlog segmentation: the table that changes the conversation
The highest-return change I recommend to any team with a backlog problem is to stop reporting one number and start reporting five. Every line is either ready to schedule or blocked by something specific, and the blocker determines who can clear it. Segmenting by blocker turns an argument about maintenance performance into a list of actions owned by named people, several of whom do not work in maintenance.
| Segment | Definition | Who clears it | What the number tells you |
|---|---|---|---|
| Ready | Approved, scoped, estimated, parts on hand, access available. Could go on next week's schedule as it stands. | Planner and scheduler | Your true capacity position. The only segment that should drive the backlog-weeks headline. |
| Waiting parts | Cannot start until a specific material arrives. Must carry the part reference and expected date. | Storekeeper, buyer, procurement | Whether stores strategy matches maintenance demand. Large and ageing here is a supply-chain issue, not a crew issue. |
| Waiting access | Blocked by an occupied room, a live process, a tenant, a clinical area, a permit not yet issued. | Operations, tenant liaison, permit authority | How much work is gated by coordination rather than resource. Large in healthcare, education and occupied estates. |
| Waiting approval | Needs a budget, quotation, client instruction or variation order before work can proceed. | Budget holder, client, finance | Where decisions are stuck. Ageing in this segment is almost always somebody declining to say no out loud. |
| Waiting shutdown | Can only be executed in a planned outage or seasonal window. Tagged to a specific future window. | Shutdown or outage planner | Your shutdown work bank. This segment is supposed to be large and old, and it should be excluded from routine backlog trending. |
Implementing this is usually one mandatory hold-reason field with a short, closed pick list, not a customisation project. What defeats it is an open text field or a list with fifteen options, because both collapse into whatever the user picks first. Keep it to five. Most systems, from IBM Maximo and Infor EAM down to MaintainX, Limble and Fiix, support this with configuration rather than development. The core CMMS modules guide sets out where this sits alongside planning and stores, and the spare parts and MRO inventory guide covers the stores discipline that shrinks the waiting-parts segment.
The political effect of this table is the real prize. Before segmentation, a rising backlog is a maintenance failure. After segmentation it is often visibly a procurement lead-time problem, a client approval problem or an access coordination problem, with the ready segment perfectly healthy. That single change moves a monthly review from defensive to productive, because the people who can unblock the work are finally looking at their own numbers.
5. Ageing: old work orders are decisions you have not made
Every backlog contains work that will never be done. A defect on a chiller replaced last year. A request from a tenant who moved out. A job raised twice. A repair a capital project has since made irrelevant. These inflate your numbers, make the report untrustworthy, and train everyone to ignore it.
Report backlog by age band, not just volume. A four-band split is enough: under 30 days, 30 to 90, 90 to 180, over 180. Then hold the honest position: anything over 180 days not tagged to a shutdown window is not backlog, it is an undecided decision. The job is either still needed, in which case it needs a priority and a date, or it is not, in which case it needs cancelling with a documented reason.
Teams avoid cancelling because it feels like an admission of failure, and in a contractual environment it can carry consequences. The reframe that works: cancelling work that is no longer needed is data hygiene, and it is materially different from failing to do needed work. Write a short cancellation reason list, and make cancellation a supervisor decision with an audit trail rather than a quiet deletion.
One caution on priority ageing. In most CMMS configurations priority is set once when the job is raised and never revisited, so a job raised as low priority nine months ago still reads as low priority even though the asset has degraded since. Any backlog review worth the name revisits priority on ageing work, particularly on critical assets. The asset criticality classification framework makes that re-prioritisation defensible rather than arbitrary, and the SLA matrix design guide covers how contractual response obligations interact with backlog priority.
6. The weekly review that actually clears backlog
Backlog does not reduce because it is reported. It reduces because somebody sits down every week and makes decisions about specific jobs. The meeting that works is short, disciplined and attended by people with authority. A structure I would recommend:
- Attendees: maintenance planner (chairs), maintenance supervisor per trade, stores or procurement representative, operations or tenant-liaison representative. Forty-five minutes, same slot every week. Without stores and operations in the room the blocked segments never move.
- Open with the one-page dashboard: backlog weeks by trade, the five segments with hours and counts, age bands, and the movement since last week. Two minutes, not twenty.
- Review the ready segment for schedulability: confirm next week's schedule is fully loaded from ready backlog, and that nothing ready is ageing past 90 days without explanation.
- Walk each blocked segment by exception: not every line, only those where the expected date has passed or the age band has moved. Every one gets a named owner and a date, recorded in the work order, not in meeting minutes.
- Take cancellation decisions explicitly: present a short list of candidates over 180 days. Cancel or re-prioritise each one in the meeting. Do not defer the decision, because deferring is how the list got to 180 days.
- Close with the shutdown work bank: confirm anything requiring an outage is tagged to a specific window with a realistic scope.
Two habits separate reviews that work from reviews that become a standing complaint session. First, decisions go into the CMMS during or immediately after the meeting, because a decision recorded only in minutes has not been made. Second, the dashboard comes from the system with no manual adjustment. The moment someone hand-edits the backlog spreadsheet before the meeting, the trend stops being real. If you are choosing a system, producing that dashboard without manual work is a genuine selection criterion; the work order software guide and the CMMS buyer's introduction both cover that ground.
The buffer also needs protecting. If urgent work raids the ready backlog every day without a decision, the planner's week is destroyed and schedule compliance collapses. That relationship is covered in the PM KPIs and schedule compliance guide, the right companion to this one.
7. Downtime: planned versus unplanned, and why the split is contested
Downtime tracking looks simpler than backlog and is usually done worse, because backlog lives entirely inside maintenance while downtime sits on the boundary between maintenance and operations, and boundaries are where data quality goes to die.
Start with the split everyone claims to have and few define properly. Planned downtime is an outage scheduled in advance with sufficient notice for operations to plan around it. Unplanned downtime is everything else. The word carrying the whole definition is "notice", and it needs a number. If a shutdown agreed at 8am for 2pm the same day counts as planned, your planned percentage will be excellent and meaningless. Pick a threshold that matches your operation, write it into the definition, and apply it without exception. Further distinctions worth making explicit before you collect anything:
- Downtime versus standby. An asset that is available but not required is not down. Redundant equipment in a duty/standby pair spends most of its life idle and healthy. If your log counts idle as down, your availability figures are nonsense.
- Downtime versus derated running. A chiller at 60 percent capacity is not down, but it is not fine either. Partial capacity loss needs a separate event type or a capacity field, or it gets logged as zero or full downtime and both are wrong.
- Downtime versus maintenance duration. The time a technician spent on the job is not the time the asset was unavailable. An asset can be down eighteen hours of which two were wrench time and sixteen were a parts wait. Both matter and they are different fields. SLA response and rectification clocks are a third thing again, with their own start and stop rules; keep them separate rather than making one field serve both.
The definitions of availability, mean time between failures and mean time to repair that consume this downtime data belong in their own discussion, and I have set them out in the reliability metrics guide. The point to carry here is simply that every one of those metrics is only as good as the downtime log feeding it, which is why the collection discipline below matters more than the formulae.
8. The downtime reason-code hierarchy
A downtime log without reason codes tells you how much time you lost and nothing about why, which supports complaint but not improvement. The fix is a reason-code hierarchy: a small top level anyone can apply correctly in seconds, and a second level carrying the analytical detail. Two levels is the right depth. One is too coarse to act on. Three means technicians guessing at the third, and a level populated by guesswork is worse than no level, because it looks like data.
| Level 1 category | Level 2 examples | Planned? | Owner of the improvement action |
|---|---|---|---|
| Equipment failure | Mechanical, electrical, instrumentation and control, structural, software or firmware | Unplanned | Maintenance and reliability engineering |
| Planned maintenance | Preventive routine, statutory inspection, overhaul, planned corrective, shutdown work | Planned | Maintenance planning |
| Maintenance delay | Awaiting spare part, awaiting specialist contractor, awaiting tool or lifting equipment, awaiting diagnosis | Unplanned | Stores, procurement, contract management |
| Operational cause | Operator error, incorrect operating regime, no demand, changeover, no operator available | Either | Operations |
| Utility or external | Power supply interruption, water supply, gas, network outage, extreme weather | Unplanned | Facilities, utility provider, business continuity |
| Safety or regulatory | Permit withheld, safety stand-down, regulatory hold, incident investigation | Either | HSE and compliance |
| Project or modification | Capital works tie-in, upgrade, commissioning, decommissioning | Planned | Projects |
The "maintenance delay" category is separated from "equipment failure" on purpose. Lumping a sixteen-hour parts wait into equipment failure makes the asset look unreliable when the real problem is the storeroom, and removes any pressure from the people who could fix it. Separating delay from failure is the most useful single change you can make to a downtime taxonomy, and the one most likely to be resisted, because it moves numbers onto other departments' reports.
Keep the list stable. A reason-code list that changes every quarter destroys year-on-year comparison, which is most of the value. Add children rarely, retire them by deactivating rather than deleting so history stays readable, and never renumber. If you are building a framework from scratch, review how failure and maintenance taxonomies are structured in the standards world, for example through ISO and the reliability-data work associated with SINTEF , which maintains OREDA. Then cut it down hard to something your own technicians will use correctly.
9. Who records downtime, and when
This is where most downtime programs quietly fail, because the data is only as good as the moment of capture and that moment is almost always inconvenient. The principle I would hold to: downtime start is recorded by whoever notices the asset stop, and downtime end by whoever returns it to service. In a plant with an operations team that means operations owns the clock and maintenance owns the reason code. In a facilities environment with no operator on the asset, maintenance owns both, and the honest consequence is that start times will often be estimates.
- Record at the event, not at close-out. Downtime entered when the work order closes three days later is reconstructed from memory and rounds to convenient numbers. A log dominated by 1, 2, 4 and 8 hour values holds estimates, not measurements.
- Make the field mandatory on the right work order types only. Mandatory downtime on every job, including a light-bulb change, teaches people to enter zero to clear the validation. Require it on corrective and emergency types against downtime-relevant assets, and nowhere else.
- Use the automated source where one exists. If SCADA, a BMS or a historian already logs asset state, take start and stop times from there and let the human supply only the reason code. Machine times plus human reasons is the strongest practical combination.
- Give mobile users a two-tap path. Capture that requires three screens on a phone in a plant room will not happen. Test this in a CMMS demonstration.
- Validate weekly, not annually. Someone should look at the previous week's downtime entries for missing reasons, overlapping events, and durations that are obviously wrong. Errors found within a week can be corrected from memory. Errors found at year end cannot.
The honest limitation on downtime data
Manually captured downtime will never be precise, and pretending otherwise is how these programs lose credibility. Start times drift, short stops go unrecorded entirely, and night-shift entries are thinner than day-shift entries on every site I have seen. The realistic goal is not precision but consistency: a log that is wrong in the same direction every month still shows you a real trend and still ranks your worst assets correctly. Report it as a trend and a ranking, never as an exact hour count in a contractual or financial context unless the times come from an automated source.
10. Asset downtime is not production downtime
Here is the disagreement that appears in almost every organisation that starts reporting downtime seriously. Maintenance reports 40 hours of asset downtime for the month. Operations says the real figure was 12. Both are reading their own records and both are right, because they are measuring different things. Asset downtime is time a specific asset was unavailable to perform its function. Production or service downtime is time the output was interrupted. The gap comes from several places, every one legitimate:
- Redundancy. One of three chillers down for eight hours is eight hours of asset downtime and zero hours of service impact, because the other two carried the load. Most of the gap on a well-designed site is this.
- Buffer and storage. A conveyor down twenty minutes with an hour of accumulation ahead of it produces no output loss.
- Non-operating hours. A machine repaired overnight while the line was not scheduled to run is asset downtime with no production consequence. Whether your downtime clock runs outside scheduled operating hours is a definition you must set explicitly, and it is the single most common cause of the disagreement.
- Cascade in the other direction. Sometimes production loss exceeds asset downtime, because a two-hour stoppage caused a four-hour restart, a reject batch or a knock-on stop downstream. Operations also logs when the line stopped while maintenance logs when the callout came in, and those are rarely the same minute.
The resolution is not to force the numbers to match. It is to publish both, labelled clearly, and to agree once with operations on the definitional questions: does the clock run outside operating hours, does redundant equipment going down count, where does the clock start. Write the answers down, put them at the front of the monthly report, and stop relitigating them. Give the service-impact figure prominence for management, because that is the one with a business consequence; the asset-level figure is the engineering number that tells you which equipment to fix.
Downtime cause data also needs to feed the schedule rather than just the report. If the reason codes show a particular asset class failing repeatedly for the same reason, that is a preventive maintenance design input. Building that loop is covered in the guide on how to build a preventive maintenance schedule.
11. Reporting both to management without inviting gaming
Every measure that carries consequence gets managed, and some of that management will be of the number rather than of the work. The way to handle it is to design the report so the easy ways to improve the number are also genuine improvements. How maintenance metrics get distorted more broadly, and the countermeasures, are set out in the PM KPIs and schedule compliance guide, so I will not repeat that ground. The distortions specific to these two measures, and the design choices that close them:
- Backlog reduced by mass cancellation. Report cancelled volume alongside completed volume, with reason. A falling backlog with a cancellation spike is then visible immediately.
- Backlog reduced by not raising work. The most damaging version, because the defects still exist and are now invisible. Trend work order creation volume and inspection-generated corrective volume on the same page as backlog. A backlog falling while defect identification also falls is a red flag, not a success.
- Backlog parked in a blocked segment. Age the blocked segments and report expected-date breaches. A waiting-parts line with no part reference and no date is a parking space, not a status.
- Downtime reduced by not recording it. Report event counts alongside hours, and track the proportion of corrective work orders on downtime-relevant assets that carry a downtime entry. A completeness measure is the only real defence.
- Downtime reclassified from unplanned to planned. Apply the notice threshold mechanically from timestamps rather than from a user-selected flag, wherever the system allows it.
- Downtime attributed to another department. Healthy if the attribution is accurate, which is why the hierarchy names an owner per category. The defence is a review of the coding, not a narrowing of the categories.
On presentation: a one-page monthly view beats a dashboard nobody opens. Backlog weeks by trade with a twelve-month trend, the five segments as a stacked bar, age bands, downtime hours split planned and unplanned with the top five contributing assets, and the level 1 reason-code distribution. Five charts, one page, definitions printed at the bottom. Where these sit in the wider performance picture is covered in the FM KPI framework, and for how they roll up into board-level measures, building a KPI tree.
One further protection is worth the effort: never set a numerical target on backlog weeks in isolation. A target of "two weeks" invites every distortion above. Set the expectation as a band with the segmentation and creation-volume context attached, and hold the team accountable for the quality of the decisions rather than the position of a single number.
The idea to walk away with
Backlog and downtime are not reporting problems, they are definition problems wearing a reporting costume. Nearly everything that makes these numbers useless is decided before any data is collected: what counts as backlog, what unit it uses, what blocks each line and who owns that blocker, what notice makes an outage planned, who starts the clock, and whether redundancy counts.
Narrow the definition to ready work, express it in crew-weeks per trade against net capacity, segment the rest by blocker with a named owner, age everything, and hold a weekly meeting where cancellation decisions are actually taken. For downtime, define the planned threshold numerically, separate maintenance delay from equipment failure, capture at the event rather than at close-out, and publish asset and service downtime as two labelled numbers rather than fighting over one. None of that requires new software. All of it requires somebody to write the definitions down and defend them for a year.
Final thoughts
The teams I have seen do this well were not the ones with the most sophisticated platform. They were the ones where the planner could tell you in one sentence what was in the backlog and what was not, and where a technician could name the five hold reasons without looking. Consistency beats sophistication in both disciplines, because the value is almost entirely in the trend and the trend only exists if the definition held still.
If you are starting from a backlog report that counts everything open and a downtime field that is half empty, do not attempt both at once. Fix the backlog definition and segmentation first, because it sits entirely within maintenance's control and produces a visible result within a month. Then take on downtime, which needs an agreement with operations and therefore needs more patience.
Disclosure
Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.
Backlog out of control, or downtime data nobody trusts?
Independent advisory on backlog definition and segmentation, planning and scheduling discipline, downtime reason-code frameworks and the reporting that stands up to a management review. 22+ years across CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.
Book a conversationRelated reading: PM KPIs and schedule compliance, Reliability metrics: MTBF, MTTR and availability, Work order types in a CMMS, Spare parts and MRO inventory in a CMMS, FM KPI framework, Building a KPI tree.
Muhammad Abbas
CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.
Work with me