Most preventive maintenance advice is written for buildings. Chillers, air handling units, fire systems, lifts: assets that sit in a facility, serve occupants, and can usually be taken offline at night or over a weekend without anyone losing money. Production plant does not behave like that. A conveyor on the main line, a screw compressor feeding the whole pneumatic ring main, a gearbox on the bottleneck machine: taking any of those offline costs output, and output is the number the plant is measured on. That single difference reshapes everything about how you plan, schedule and justify industrial preventive maintenance, and it is the reason plant maintenance planners end up building their own schedules rather than using generic facilities templates.
The message up front: in a production environment PM competes directly with output, so the schedule is a negotiation, not a decree. Build the programme around three things: criticality ranked by position in the production flow rather than by replacement cost, a planned shutdown rhythm that absorbs the intrusive work, and operator care that catches the cheap problems daily. A spreadsheet will carry that programme honestly up to roughly a few hundred assets and one site. Past that, it quietly starts lying to you.
1. What is genuinely different about plant PM
Before the schedules and the tables, it is worth being precise about the differences, because they are not cosmetic. If you carry facilities habits into a plant, the programme will be rejected by operations within a quarter.
- PM consumes production capacity. In a building, a two-hour AHU service costs two hours of a technician. On a production line, a two-hour intervention costs two hours of technician plus whatever the line produces in two hours, plus ramp-down and ramp-up losses either side. The true cost of a PM task in a plant is often ten to fifty times the labour cost, and until you price it that way you will not understand why operations resist your schedule.
- Access is negotiated, not booked. A facilities planner books a window. A plant planner requests one, and gets it when the production plan allows. The planner who understands the production schedule, the changeover points, the order book and the seasonal demand curve gets far more access than the planner who simply issues work orders and complains they are not done.
- Failure consequence is measured in throughput, not comfort. A failed lighting circuit in an office is an inconvenience. A failed drive on the bottleneck machine is lost revenue for every minute it is down, and possibly scrapped work in progress as well. Criticality has to be built on this logic, which I will come back to in detail.
- Operators are part of the maintenance system. In facilities, the occupant is a passive user. In a plant, the operator stands at the machine for eight or twelve hours a day, hears it, feels it and knows its moods better than any technician on a monthly round. Ignoring that is throwing away your best early-warning sensor.
- The dominant rhythm is the shutdown. Facilities PM spreads work evenly through the year. Plant PM concentrates the intrusive work into planned shutdowns and turnarounds, and fills the rest of the year with what can be done while running. The annual plan is shaped around those windows.
If you are coming to this from the general discipline, the foundations still hold: the strategy taxonomy in the complete guide to preventive maintenance and the frequency logic in preventive maintenance strategies apply just as much to a press line as to a chiller plant. What changes is the scheduling constraint and the economics around it.
2. OEE: the number the plant is actually judged on
Overall Equipment Effectiveness is the language production speaks, and a maintenance manager who cannot speak it will lose every scheduling argument. OEE multiplies three loss categories:
Availability = run time / planned production time
lost to: breakdowns, changeovers, waiting for materials
Performance = (ideal cycle time x units) / run time
lost to: minor stops, reduced speed, jams
Quality = good units / total units
lost to: scrap, rework, startup rejects
Maintenance people instinctively reach for availability, because breakdowns are the visible loss. But in my experience the more persuasive maintenance argument usually sits in performance and quality. A worn conveyor belt that causes fifteen micro-stops a shift never appears as a breakdown, never generates a corrective work order, and can quietly cost more output than the one dramatic failure everyone remembers. A drifting instrument that pushes a process to the edge of tolerance shows up as scrap on the quality line, not as a maintenance number at all.
The practical move is to ask production for their minor-stop and scrap data by machine, and look for the maintenance causes hiding inside it. Chronic jams, speed reductions the operators have normalised, startup rejects after every changeover: these are maintenance defects wearing a production costume. Presenting a PM proposal in terms of recovered performance loss, rather than hypothetical avoided breakdowns, is what gets you the shutdown window.
The argument that wins the window
Do not ask operations for four hours to do maintenance. Show them the machine losing eleven minutes a shift to minor stops, tell them which PM task removes that cause, and ask for four hours to recover fifty-five minutes a week. The conversation changes from cost to return, and it is a conversation the production manager can take to their own boss.
3. Criticality driven by bottleneck position, not replacement cost
The standard criticality method ranks assets by consequence of failure across safety, environment, production and cost. That is sound, and the general method is covered in asset criticality classification. But in a plant it needs one specific adjustment that generic models get wrong: position in the production flow matters more than the value of the asset.
A cheap transfer conveyor sitting between two expensive machines, with no bypass and no buffer either side, can stop the whole line. An expensive spare press running in parallel with two others, at sixty percent loaded, can fail on a Friday and nobody notices until Monday. Rank those by replacement cost and you will maintain the wrong one. The questions that actually determine plant criticality:
- Is it the constraint? Every line has a bottleneck. An hour lost at the bottleneck is an hour lost by the whole plant and can never be recovered. An hour lost at a non-bottleneck is usually absorbed by slack. Identify the constraint machine explicitly and rank it above everything else.
- Is there redundancy or a bypass? A duty/standby pump pair is materially less critical than a single pump doing the same job, provided the standby is actually tested and actually works.
- Is there buffer either side? Work in progress between stations decouples them. A machine with thirty minutes of buffer downstream can be stopped for twenty minutes with no line impact. A machine with no buffer cannot.
- What is the restart cost? Some processes restart in seconds. Others, particularly thermal, chemical and coating processes, need hours to stabilise and scrap everything made in between. Restart cost can dwarf the downtime itself.
- What is the spare lead time? An asset whose critical part takes sixteen weeks to source is functionally more critical than an identical asset with a part on the shelf, because the failure consequence is measured in months not hours.
Here is a scoring model I would recommend building, sized so it fits in a spreadsheet and can be argued through with a production manager in one workshop. Score each factor 1 to 5, weight, sum.
| Factor | Weight | Score 1 (low) | Score 3 (medium) | Score 5 (high) |
|---|---|---|---|---|
| Safety / environmental | x5 | No credible harm | Injury or reportable release possible | Major injury, fire or regulatory breach credible |
| Bottleneck position | x5 | Off-line, parallel, spare capacity | On line, not the constraint | The constraint machine |
| Redundancy | x4 | Full standby, regularly proven | Partial or manual bypass | Single point of failure |
| Buffer downstream | x3 | Hours of work in progress | Minutes of buffer | Direct coupling, no buffer |
| Restart / stabilisation cost | x3 | Restart in minutes, no scrap | Under an hour, some scrap | Hours to stabilise, batch lost |
| Spare lead time | x3 | On shelf | 2 to 6 weeks | Over 12 weeks or obsolete |
| Failure history | x2 | No failures in 3 years | Occasional, predictable | Repeat offender, unpredictable |
| Banding: 100+ = Critical (A), full PM plus condition monitoring, spares held · 60 to 99 = Important (B), standard PM, spares sourced on plan · below 60 = Routine (C), minimal PM or run to failure with a fast corrective response | ||||
Two disciplines make this stick. First, run the scoring with production in the room, not in the maintenance office, because the bottleneck and buffer answers are theirs and the agreement is worth more than the score. Second, re-run it when the product mix changes. The bottleneck moves when the order book moves, and a criticality ranking built around last year's product mix will be quietly wrong.
4. The equipment classes and what each actually needs
Plant PM content is driven by failure mode, not by a generic template. These are the classes that make up most of an industrial register and the tasks that genuinely earn their place on each.
Conveyors and material handling. The most-neglected high-impact asset class in most plants, because conveyors look simple and are cheap individually. The failure modes that matter are belt tracking and wear, roller and idler seizure, drive and gearbox condition, take-up tension, and guard and pull-cord integrity. Weekly walk-the-line inspections looking for seized rollers by touch or thermal gun, monthly tension and tracking checks, quarterly drive inspection. Seized idlers are worth hunting specifically: they burn belts, they are cheap to replace, and they are invisible until the belt is scrapped.
Compressors and compressed air systems. Compressed air is the most expensive utility in most plants and the one with the least oversight. PM content: filter and separator changes on hours, oil analysis and change, condensate drain function, intercooler and aftercooler cleaning, dryer dewpoint verification, and above all leak surveys. Ultrasonic leak detection on the ring main is one of the highest-return maintenance activities available in a factory, and it does not need a shutdown. Treat the distribution network as an asset, not just the compressor.
Hydraulics. Contamination is the dominant failure cause, so fluid cleanliness is the PM. Sample and test oil to a cleanliness standard, change filters on differential pressure rather than blind calendar, check reservoir level and temperature, inspect hoses for abrasion and age, verify accumulator pre-charge, and log relief and setting pressures so drift is visible. A hydraulic system that is kept clean and cool outlives one that is rebuilt regularly but run dirty.
Gearboxes and drives. Oil analysis is the single most informative task, giving the longest warning of any technique through wear metals and viscosity change. Add vibration on the critical units, breather condition (a blocked breather draws moisture in on every thermal cycle), coupling alignment after any intervention, and belt tension and sheave wear on belt drives. Thermal checks on motor bearings on the route catch the rest.
Process pumps and valves. Seal condition, bearing vibration and temperature, coupling alignment, suction strainer condition, and performance drift: a pump whose discharge pressure is falling at constant speed is telling you something. For valves, stroke and exercise the ones that rarely move, because the emergency isolation valve that has not been operated in three years is the one that will not close when you need it. Actuator and positioner calibration belongs on the instrumentation schedule.
Instrumentation and calibration. The class that quietly drives quality loss. Loop calibration of temperature, pressure, flow and level instruments on a documented interval with traceability; weighing and filling equipment calibration; safety instrumented function proof testing on its own interval with its own records. This is also where audit and regulatory exposure concentrates, so keep the certificates with the work order. The ISO 9001 measurement-traceability requirements make this non-optional in most manufacturing settings; the standard itself is described at ISO .
Electrical distribution and control panels. Thermographic survey of panels and switchgear under load (annually at minimum), torque checks on terminations, cleanliness and filter condition on panel cooling, and drive and inverter fan and capacitor checks. Panel cooling failure is a slow killer of drives that nobody notices until the fault codes start.
Where I would not add PM
Resist the urge to write a PM for everything in the class. Intrusive maintenance on a healthy machine introduces its own failures: disturbed terminations, contaminated hydraulics, misaligned couplings, wrongly reassembled guards. On C-band assets with cheap, fast-replaceable failure modes, a fast corrective response and a shelf spare beat a quarterly inspection that nobody has time to do properly. A shorter programme executed at ninety-five percent compliance outperforms a comprehensive one executed at forty.
5. Autonomous maintenance: the operator as the first line
The Total Productive Maintenance idea that translates best into ordinary plants is autonomous maintenance: giving operators a small, well-defined set of daily care tasks on the machine they run. Not repairs, and not anything requiring isolation. Clean, inspect, lubricate, tighten, report.
The value is early detection. A technician on a monthly route sees the machine twelve times a year. The operator sees it two hundred and forty times. Loose guards, small leaks, unusual noise, a hot bearing, a build-up of product in the wrong place: the operator will catch all of these weeks before the route does, if you give them a structure and a route for reporting that actually leads somewhere.
What makes an autonomous maintenance programme work, in the plants where I have seen it stick:
- Keep the task list very short. Five to ten items, doable in the first ten minutes of the shift. A thirty-item operator checklist becomes a pencil-whipped ritual within a month.
- Make it visual. Marked gauge ranges, direction-of-rotation arrows, oil level sight glasses with min and max painted on, colour-coded lubrication points. If the operator has to read a value and remember the limit, they will not. If green is good and red is not, they will.
- Close the loop on reports. Nothing kills operator reporting faster than raising three defects and hearing nothing. Every operator-raised defect should become a visible work request with a visible outcome. This is the most common point of failure and it is entirely a process problem, not a people problem.
- Train on the why, not just the what. An operator who understands that the greasing point they are hitting feeds a bearing that costs the line four hours when it seizes will do it properly. One who was handed a checklist will not.
- Do not use it to absorb headcount cuts. The moment autonomous maintenance is perceived as maintenance work pushed onto operators to justify reducing the maintenance team, it is dead. Frame it as early detection, and keep the technical work with technicians.
6. Planned shutdowns and turnarounds: the dominant rhythm
In a continuous or heavily loaded plant, the intrusive PM work does not happen weekly. It happens in the shutdown, and the annual maintenance plan is really a plan for how to fill those windows. Getting shutdown planning right is worth more than any individual PM task.
Building the scope. A shutdown scope should be assembled from five sources, each with a named owner:
- Scheduled PM falling due that cannot be done while running: internal inspections, alignment, anything needing isolation or entry.
- Deferred corrective work from the backlog, filtered to what genuinely needs the outage. Backlog review is the single richest source and the most neglected.
- Condition-monitoring findings: the bearing trending up, the thermography exception, the oil sample with rising iron. This is where condition monitoring and failure prediction pays for itself, by converting a future breakdown into a shutdown line item.
- Statutory and insurance inspections: pressure systems, lifting equipment, safety valves. These have immovable dates and should anchor the shutdown calendar rather than be fitted around it.
- Projects and modifications: capital work, upgrades, obsolescence replacement. These compete for the same window and the same isolations, so they belong in the same scope list.
Freeze the scope. The discipline that separates a shutdown that finishes on time from one that overruns is a scope freeze date, typically six to eight weeks before for a routine shutdown and considerably longer for a major turnaround. After the freeze, additions require a named approver and a stated impact on the critical path. Without this, scope creep will consume the contingency before the shutdown starts.
Plan the long-lead items backwards from the freeze. Materials, specialist contractors, cranes, scaffolding, isolation and permit preparation all have lead times. Working backwards from the shutdown date through the freeze date to the order dates is the planning sequence. Discovering at the freeze that a critical part has a fourteen-week lead is how shutdowns get postponed.
The shutdown readiness test
Two weeks out, ask one question of every job in the scope: are the parts on site, the permits drafted, the isolations identified, the labour named and the method agreed? Any job that cannot answer yes to all five is not ready and should either be recovered urgently or removed from the scope. Jobs carried into a shutdown unprepared are the ones that overrun and take the rest of the plan with them.
7. A plant PM schedule structure that works in a spreadsheet
Most plant maintenance programmes start life as an Excel schedule, and for a single site of moderate size that is a perfectly respectable place to start. What matters is the structure. Build it as a flat, one-row-per-task master list, not as a wall calendar, because a flat list can be sorted, filtered, pivoted and eventually imported into a CMMS. A calendar-shaped spreadsheet cannot.
The columns I would insist on:
| Column | Example | Why it is there |
|---|---|---|
| Asset ID | LN2-CNV-014 | Unique, structured tag. The join key to everything else. |
| Asset description | Line 2 infeed conveyor | Human readable, for the technician. |
| Line / area | Line 2 · Packing hall | Lets you filter the schedule by the production area you have access to. |
| Equipment class | Conveyor | Drives task content and lets you audit coverage by class. |
| Criticality band | A | From the scoring table. Drives frequency and spares policy. |
| Bottleneck flag | Yes | The constraint filter. Answers "what must never stop". |
| PM task ID | CNV-M-01 | Reusable task code. Same code across every conveyor. |
| Task description | Belt tracking and tension check | What is actually done. |
| Frequency | Monthly | Or a meter value. Keep the vocabulary controlled. |
| Basis | Calendar | Calendar, running hours, cycles, or condition. |
| Running / shutdown | Running | The single most important plant-specific column. Determines whether it needs a window. |
| Duration (hrs) | 0.5 | Feeds the labour load calculation and the shutdown critical path. |
| Trade / crew | Mechanical | For levelling the workload by discipline. |
| Isolation required | LOTO, mechanical | Flags permit and safety preparation. Never leave this implicit. |
| Parts / consumables | Belt cleaner blade | Feeds the kitting list and the min/max stock review. |
| Last done | 2026-08-14 | The anchor for the next due date. |
| Next due | 2026-09-14 | Calculated, not typed. This is where spreadsheets get it wrong. |
| Status | Due | Due, overdue, done, deferred. Drives the compliance figure. |
| Owner | Shift 2 mech | Named accountability. Unowned tasks do not happen. |
| Reference doc | OEM man. s.6.2 | Traceability back to manufacturer or standard. |
A few practical rules for making this work in Excel. Put the master task list on one sheet and nothing else on it: no merged cells, no colour-as-data, no blank separator rows. Drive Next due with a formula from Last done and Frequency, never by typing dates, because hand-typed due dates drift within two months. Use a pivot table on Trade and Month to see labour load and spot the weeks where you have promised forty hours of work to a crew of two. Filter on Running/Shutdown to produce the shutdown scope list directly. And keep a separate flat sheet as a completion log, one row per completion with date, who, and findings, rather than overwriting the Last done cell, because the overwrite destroys the history you will need later.
For task content and checklist wording, rather than reinventing every line, start from the structures in PM schedule and checklist templates and the build sequence in how to build a preventive maintenance schedule, then adapt the frequencies to your duty and environment.
8. A sample plant PM schedule you can lift into Excel
This is a realistic skeleton across the main industrial classes, at frequencies that suit a two-shift operation in a moderately dusty environment. Treat the frequencies as a starting position to be adjusted by duty, environment and failure history, not as gospel. OEM manuals and your own failure data outrank any published table, including this one.
| Class | Task | Freq | Basis | Run / SD | Trade |
|---|---|---|---|---|---|
| Conveyor | Operator check: guards, noise, spillage, e-stop visual | Daily | Calendar | Run | Operator |
| Conveyor | Idler and roller condition survey, thermal scan for seizure | Weekly | Calendar | Run | Mech |
| Conveyor | Belt tracking, tension, cleaner blade, splice inspection | Monthly | Calendar | Run | Mech |
| Conveyor | Drive gearbox oil sample, coupling alignment check | 6 months | Calendar | SD | Mech |
| Compressor | Operator check: pressure, temp, condensate drain function | Daily | Calendar | Run | Operator |
| Compressor | Intake and panel filter condition, cooler fin cleanliness | Monthly | Calendar | Run | Mech |
| Compressor | Oil and separator change, oil analysis | 2,000 hrs | Running hrs | SD | Mech |
| Compressed air | Ultrasonic leak survey of ring main and drops | Quarterly | Calendar | Run | Mech |
| Compressed air | Dryer dewpoint verification, drain trap test | Quarterly | Calendar | Run | Mech |
| Hydraulics | Level, temperature, visible leak and hose abrasion check | Weekly | Calendar | Run | Operator |
| Hydraulics | Fluid sample for cleanliness and water content | Quarterly | Calendar | Run | Mech |
| Hydraulics | Filter change on differential pressure, accumulator pre-charge | As indicated | Condition | SD | Mech |
| Gearbox / drive | Vibration reading on critical units | Monthly | Calendar | Run | Reliability |
| Gearbox / drive | Oil sample, breather condition, external leak check | Quarterly | Calendar | Run | Mech |
| Gearbox / drive | Oil change, internal inspection where access allows | Annual | Calendar | SD | Mech |
| Belt drive | Tension, sheave wear, guard integrity | Monthly | Calendar | SD | Mech |
| Process pump | Vibration and bearing temperature, seal leak check | Monthly | Calendar | Run | Mech |
| Process pump | Performance check: flow and discharge pressure vs curve | Quarterly | Calendar | Run | Mech |
| Process pump | Coupling alignment verification, strainer clean | Annual | Calendar | SD | Mech |
| Valves | Exercise and stroke rarely-operated isolation valves | 6 months | Calendar | SD | Mech |
| Valves | Safety relief valve test and certification | Statutory | Calendar | SD | Specialist |
| Instrumentation | Loop calibration: critical process temp, press, flow, level | 6 months | Calendar | SD | Instrument |
| Instrumentation | Weighing and filling equipment calibration | Quarterly | Calendar | Run | Instrument |
| Instrumentation | Safety instrumented function proof test | Per SIL study | Calendar | SD | Instrument |
| Electrical | Thermographic survey of panels and switchgear under load | Annual | Calendar | Run | Elec |
| Electrical | Panel cooling filter clean, fan function check | Quarterly | Calendar | Run | Elec |
| Electrical | Termination torque check, drive capacitor and fan inspection | Annual | Calendar | SD | Elec |
| Material handling | Forklift and lifting equipment statutory examination | Statutory | Calendar | Run | Specialist |
| Material handling | Chain, sling and hoist visual inspection | Monthly | Calendar | Run | Mech |
Note the split in the Run/SD column. Roughly two thirds of the tasks here can be done while the plant runs. That ratio is the target: the more inspection and monitoring you can move off the shutdown list, the less you compete with production and the shorter your shutdowns become. Every task you can convert from shutdown to running is a task you will actually complete.
9. When the spreadsheet stops working (be honest about this)
I am not going to pretend Excel is a maintenance system. It is an excellent design tool and a serviceable execution tool for a while, and a great many plants run a competent programme this way. But there are specific, recognisable points where it stops being honest, and pushing past them costs more than the software would have.
- You cannot prove compliance. The moment an auditor or a customer asks for evidence that the June calibration was done, by whom, with what result, and the answer is a cell that says "done", the spreadsheet has failed. There is no record, no timestamp, no attachment, no signature.
- History is being overwritten. Every time Last done is updated, the previous value is gone. Without history you cannot compute MTBF, spot repeat failures, or justify a frequency change. You are maintaining without learning.
- More than one person needs to edit it. Two shifts, three planners and a contractor in one workbook produces conflicting copies, lost updates and a file named final_v7_REAL.xlsx. This usually happens around forty to sixty active users of the information.
- Corrective work has nowhere to live. PM is only part of the workload. Breakdowns, operator-raised defects and shutdown jobs need to sit alongside PM in one backlog so you can prioritise across them. A PM spreadsheet cannot hold a backlog, which is where the work order type structure starts to matter.
- Meter-based tasks are due and nobody knows. Anything triggered by running hours or cycles needs a meter reading imported from the machine or the historian. Manually transcribing hour meters into a spreadsheet works for five machines and fails for fifty.
- Spares are not connected. When you cannot see from the task what parts it consumes, and from the part what stock remains, kitting becomes guesswork and stockouts become normal.
- Multiple sites, one standard. Two plants running two spreadsheets will diverge within a year. There is no way to enforce a common task library or compare performance across sites.
My honest rule of thumb: a single site, under roughly three to five hundred assets, one or two planners, no regulatory evidence burden, and no meter-driven tasks is a reasonable spreadsheet case. Cross two or three of those lines and the spreadsheet is costing you more in lost history, rework and unprovable compliance than a modest system would cost to run. If that is where you are, the selection logic is in the CMMS buyer shortlist. The one thing I would say plainly: design the programme properly in the spreadsheet first. A well-structured spreadsheet imports cleanly into any CMMS. A badly structured one carries its mess into the new system and you will pay for that for years.
The software will not fix the programme
Moving a poorly designed PM programme from Excel into IBM Maximo, SAP PM, Infor EAM, Hexagon EAM, Limble or MaintainX does not improve it. It makes the same wrong frequencies easier to generate, the same unnecessary tasks harder to delete, and the same compliance gap more visible. The system solves record-keeping, scheduling and evidence problems. It does not solve a programme that was never matched to failure modes in the first place.
10. Spares strategy, especially long-lead items
In a plant, the spares policy is part of the maintenance strategy, not a stores problem to be handled separately. The deciding variable is lead time against the cost of being down for that lead time.
The classification I would recommend:
- Insurance spares. High value, very long lead, catastrophic if unavailable. A main drive motor, a large gearbox, a critical PLC rack, a specialist rotor. You hold one not because it will be consumed but because the plant cannot survive the lead time. Justify these individually with a written case tied to the criticality score, and review them annually. They are expensive, and the temptation to release the capital in a cost-cutting year is exactly the decision that hurts three years later.
- Critical consumables. Bearings, seals, belts, filters, sensors on A-band assets. Moderate cost, predictable consumption. Min/max stock driven by usage history and lead time. These should be reviewed against actual consumption every six months, because PM changes move consumption.
- Kitted PM parts. Parts consumed by scheduled tasks. These should be picked and staged against the work order before the window, not fetched during it. Kitting is one of the cheapest wrench-time improvements available.
- Commodity items. Fasteners, general lubricants, standard electrical parts. Vendor-managed or consignment stock wherever possible so they do not consume planner attention.
Three things worth doing specifically about long-lead items. First, build an obsolescence register: any drive, controller or instrument whose manufacturer has issued an end-of-life notice moves to the top of the risk list, because lead time on an obsolete part is effectively infinite. Second, record the actual quoted lead time as a field on the part, refreshed annually, rather than trusting a number entered at handover. Third, when a supplier quotes a long lead, ask whether a repairable core or a refurbishment route exists, because a repair exchange programme on large rotating equipment often beats holding a new unit.
Where formal rigour is warranted on the highest-consequence assets, the criticality and spares decisions both follow naturally from an RCM analysis, which forces the question of what each failure mode actually does and what is worth holding against it.
11. Measuring an industrial PM programme
The measures that tell you whether a plant PM programme is working, and the ones I would put on a single monthly page:
- PM compliance, measured as tasks completed within their allowed window, not simply completed eventually. Track A-band compliance separately; ninety percent overall with sixty percent on critical assets is a failing programme wearing a good number.
- Planned versus unplanned work ratio, by labour hours. A plant moving toward eighty percent planned is maturing. One stuck below fifty is firefighting regardless of what the PM schedule says.
- Unplanned downtime hours on the bottleneck. Not plant-wide downtime, which averages away the thing that matters. The constraint machine specifically.
- MTBF on A-band assets, trended over twelve months. The one measure that actually says whether reliability is improving.
- Backlog size and age, in crew-weeks. A backlog of four to six crew-weeks is healthy. Zero means you are not finding defects. Twenty means you are not resourcing them.
- Schedule break-in percentage: how much of the week's planned work was displaced by emergent work. High break-in explains poor compliance better than any other number.
The broader KPI structure, including how to avoid the common trap of measuring maintenance activity rather than plant outcomes, is in the KPI framework. The plant-specific addition is to report at least one maintenance measure inside the production OEE conversation, so the two functions are looking at one scoreboard rather than two.
On statutory and safety-critical measures, keep the evidence separate and absolute. Lifting equipment examinations, pressure system inspections, safety function proof tests and electrical testing are not subject to the same prioritisation logic as everything else; they are done on date, and the certificate is the deliverable. Guidance on the machinery safety and inspection regimes in most jurisdictions traces back to the relevant standards bodies, for example the standards catalogues at ISO and the electrical safety and equipment standards at IEC .
The idea to walk away with
Industrial preventive maintenance is not facilities PM applied to bigger machines. It is a scheduling negotiation conducted in the language of production, anchored on three structures: criticality that follows the production flow rather than the asset register value, a planned shutdown rhythm that absorbs everything intrusive, and operator care that catches the cheap problems two hundred and forty times a year instead of twelve.
Build the schedule as a flat, structured task list rather than a calendar, split every task by whether it can be done running or needs an outage, and drive the due dates by formula rather than by hand. That structure works in a spreadsheet for a single moderate site, and it is exactly the structure that imports cleanly the day the spreadsheet runs out. The programme design is the durable asset; the tool it lives in is replaceable.
Final thoughts
The plants with the best maintenance performance I have seen are rarely the ones with the most sophisticated systems. They are the ones where the maintenance planner knows the production schedule as well as the production planner does, where the operators report defects because they know something happens when they do, where the shutdown scope is frozen and prepared rather than improvised, and where a short list of well-chosen tasks gets done at high compliance instead of a long list getting done at forty percent.
If you are starting from nothing, do not start by buying software and do not start by writing four hundred PM tasks. Start by identifying the constraint machine, scoring thirty assets with a production manager in the room, and writing a tight PM set for the A band with the running and shutdown split marked. That fits comfortably in a spreadsheet, it can be built in a fortnight, and it will produce more reliability than a comprehensive programme that nobody has the hours to execute. Everything else can be added once the core is proven. For the programme-level structure that this eventually grows into, see the PM plans and programmes framework.
Building or fixing a plant PM programme?
Independent advisory on criticality ranking, PM programme design, shutdown planning and CMMS/EAM selection and implementation. 22+ years across manufacturing, utilities, oil and gas, government and facility operations. No vendor margins, no reseller arrangements.
Book a conversationRelated reading: Preventive maintenance: the complete guide, How to build a preventive maintenance schedule, PM schedule and checklist templates, Asset criticality classification, Predictive maintenance and failure prediction, CMMS software buyer shortlist.
Muhammad Abbas
CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.
Work with me