Most facilities run maintenance to protect asset life and control cost. A data centre runs maintenance to protect uptime, and the uncomfortable arithmetic behind the industry is that a large share of critical facility outages are caused not by equipment failing on its own but by people working on equipment. Wrong breaker, wrong sequence, wrong valve, right procedure but the out-of-date revision. That single fact reorders everything a CMMS has to do in this sector. Scheduling, approvals, document control, evidence capture and vendor management all stop being back-office hygiene and become the mechanism by which risk is either contained or created. This guide is about building that mechanism properly.
The message up front: in a critical facility the CMMS is not primarily a work scheduler, it is a change-control and evidence system that happens to schedule work. If it cannot prove which procedure revision was used, who approved the intervention, which redundant path was isolated and what the environmental conditions were while the work happened, it is not fit for a data centre no matter how good its mobile app is.
1. Why data centre maintenance is a different discipline
A commercial office tower can tolerate a chiller offline for an afternoon. Occupants complain, nobody breaches a contract. A colocation hall cannot tolerate the same event unless the redundant path is proven healthy first, because the operator has signed availability commitments with financial teeth and, more importantly, because customers have built their own commitments on top of those. The consequence profile is what makes this a distinct discipline, and it shows up in five concrete ways.
- The work itself is the hazard. In most sectors the failure is the risk and maintenance is the mitigation. In a critical facility maintenance is also a risk, because it involves touching live infrastructure that thousands of workloads depend on. This inverts the usual instinct to do more work more often.
- There is no natural downtime window. No overnight shutdown, no seasonal closure, no quiet weekend. Every intervention happens on a facility that is running, so every intervention needs a designed safe state rather than a convenient hour.
- The asset population is narrow and deep. Compared with a mixed-use estate the asset classes are few: electrical distribution, UPS and batteries, generation, cooling, fire and life safety, leak detection, security. But each one is instrumented, redundant, warranty-bound and often vendor-maintained under contract, so the depth of record per asset is far greater.
- Evidence is a deliverable. Customers audit you. Certification schemes audit you. In many cases the maintenance record is contractually disclosable. The CMMS output is not internal, it faces outward.
- Vendors do much of the work. A large proportion of critical plant maintenance is delivered by the OEM or an authorised service partner, not by the in-house team. Your CMMS has to manage work you did not perform and evidence you did not generate.
If you are coming to this from a general facilities background, the foundations still apply and are worth having straight first: see CMMS for facilities management and, if the category itself is new to you, what a CMMS actually is. What follows assumes those basics and adds the critical-facility layer on top.
2. Concurrent maintainability: what redundancy means for how you schedule
The two concepts that govern data centre maintenance planning are concurrent maintainability and fault tolerance, and the distinction between them decides what you are allowed to do while the facility is live.
Concurrent maintainability means any capacity component or distribution element can be removed from service for planned maintenance without affecting the critical load. You can take a UPS module, a chiller, a pump or a distribution path out deliberately, because the design provides an alternative. Fault tolerance is a stronger property: the facility can also absorb an unplanned single failure without dropping load, including a failure that occurs while something else is already out for maintenance. The tiering frameworks published by the Uptime Institute formalise these properties, and whichever framework your site was designed to, the operational consequence is the same.
That consequence is a rule I would put on the wall of every critical facility planning office: you maintain the redundant path, never the live one. The sequence is always the same. Verify the alternative path is healthy and capable of carrying the load. Transfer or isolate so that the target equipment is no longer carrying critical load. Maintain it. Restore it. Verify. Only then consider the next item. The maintenance activity follows the redundancy, and the redundancy is what makes the activity survivable.
This has three practical implications your CMMS has to support and most generic configurations do not.
- Redundancy pairs must be modelled, not remembered. The asset record for UPS-A needs to know that UPS-B is its partner, and the system needs to be able to refuse or flag a work order that would put both out of service in the same window. If that relationship lives only in a senior engineer's head, it leaves when they do.
- Concurrency rules beat calendar efficiency. The natural planner instinct is to batch similar work: do all four CRAH units in one visit, all the battery strings in one day. In a critical facility that instinct is dangerous. Schedule by path, not by asset class, even though it costs more mobilisations.
- Reduced-redundancy time is a tracked quantity. While UPS-A is out, the site is running at reduced resilience. That window should be explicitly bounded, visible on a dashboard, counted, reported and minimised. Teams that track cumulative hours at reduced redundancy per quarter make far better decisions about scope batching than teams that do not.
The scheduling test
Before a critical work order is released, someone must be able to answer, from the system and not from memory: which redundant path carries the load during this work, has that path been verified healthy, what is the maximum time we accept at reduced redundancy, and what is the abort trigger that sends us back to the original configuration. If the CMMS cannot hold those four answers against the work order, it is not carrying its share of the risk control.
3. Every physical intervention is a change
The most important cultural import into critical facility maintenance comes from IT service management: the idea that nothing changes without an approved change record. In a data centre this extends to the physical plant. Replacing a filter is a change. Opening a valve is a change. Racking a breaker is a change. Firmware on a UPS is emphatically a change. The reason is not bureaucratic taste, it is that the facility's resilience is a designed state, and any intervention temporarily departs from that state.
A workable change model for physical maintenance has four tiers, and the value of tiering is that it stops the process collapsing under its own weight.
| Change class | Typical work | Approval path | Evidence required |
|---|---|---|---|
| Pre-approved / routine | Visual inspections, readings, filter checks, non-intrusive rounds | Standing authorisation under a named SOP | Completed round record, readings, exceptions noted |
| Standard planned | Routine PM on redundant plant: CRAH service, generator run, pump PM | Operations manager plus shift lead on the day | MOP revision used, isolation record, before/after readings |
| Critical planned | Switchgear work, UPS bypass, battery string replacement, transfer testing | Change advisory review, customer notification where SLA requires | Signed MOP, risk assessment, PTW, back-out plan, witness sign-off |
| Emergency | Corrective response to a live fault or an alarm that will escalate | Duty authority on the spot, retrospective review mandatory | EOP referenced, timeline reconstruction, post-incident record |
The CMMS role here is to be the place the change record and the work order are the same object, or at minimum are linked so tightly that one cannot be closed without the other. I have seen sites run a change register in a service-management tool and the work orders in a CMMS with no link between them, and the result is predictable: a reconciliation exercise before every audit, and a genuine inability to answer whether a given intervention was ever approved. If you must run two systems, integrate them at the point of release and the point of closure, and make the change reference a mandatory field on the critical work order.
Permit to work sits inside this structure rather than beside it. Electrical isolation, hot work, confined space and work at height permits should be raised from the work order, carry the same asset and isolation references, and be closed before the work order can be closed. The general pattern for wiring this properly is in permit to work integration with a CMMS, and in a data centre the only change is that almost nothing on the critical path is permit-exempt.
4. MOP, SOP, EOP: the document set and why versioning is the whole game
Critical facilities operate on a standardised procedure set. The naming varies slightly between operators but the structure is consistent.
- SOP, standard operating procedure: how a system is normally operated. Normal configuration, normal parameters, routine actions. The baseline description of correct operation.
- MOP, method of procedure: the step-by-step script for a specific planned intervention. Prerequisites, tools, isolation steps, the work itself, restoration steps, verification, and a back-out plan for each decision point. A MOP is written per task and per configuration, and a good one is unambiguous enough that a competent stranger could execute it.
- EOP, emergency operating procedure: what to do when something has gone wrong. Loss of utility, UPS fault, cooling failure, fire alarm activation, leak detected. Written to be followed under stress by whoever is on shift.
- SMOP or scripted MOP: some operators distinguish a fully scripted, signed-off MOP for the highest-risk activities, executed line by line with a second person reading and initialling each step.
The maintenance management requirement is narrower than the document library question and much more important: the CMMS must carry the current, approved revision of the applicable procedure against the asset, and the work order must record which revision was actually used. Those are two different things and both matter.
Carrying the current revision against the asset means a technician opening a work order on UPS-2 does not go hunting on a file share. The document, at its approved revision, is attached to or linked from the asset and inherited by the work order. Recording the revision used means that six months later, when a similar intervention goes wrong, you can establish whether the earlier job followed the same script. Without that, incident investigation becomes archaeology.
Where most CMMS platforms are genuinely weak
Document control is the honest gap. Most CMMS products, including good ones, offer file attachment rather than controlled document management: no revision workflow, no approval states, no automatic supersession, no guarantee that the attached file is the current one. Enterprise EAM platforms such as IBM Maximo and Hexagon EAM can be configured closer to real revision control, usually with effort. The pragmatic pattern I would recommend is to keep procedures in a proper document management system, reference them from the asset and work order by document number plus revision, and have the CMMS store the revision string rather than a copy of the file. It is less elegant than one system holding everything, and it is far safer than a stale PDF attached to an asset record in 2023.
5. The critical systems and what maintenance management they demand
The asset classes below are the core of any data centre maintenance programme. The point of tabulating them is not to prescribe intervals, which depend on the equipment, the manufacturer and the operating environment, but to show how differently each one behaves as a maintenance management problem.
| System | Dominant maintenance concern | Work control characteristic | CMMS demand |
|---|---|---|---|
| UPS modules | Capacitors, fans, firmware, bypass integrity | Bypass operation is a critical change; almost always OEM-delivered | Vendor work orders, firmware version history, bypass event log |
| Battery systems | String impedance trend, cell voltage, temperature, end of life | String isolation reduces autonomy; replacement is capital-planned | Per-string and per-cell measurement history, trend-based replacement forecasting |
| Generators | Start reliability, fuel quality, load acceptance, cooling | Load-bank and black-start testing are scheduled critical events | Run-hour meters, test result records, fuel sampling history |
| Switchgear & distribution | Thermographic anomalies, torque integrity, protection settings, breaker exercise | Highest-consequence work in the building; full isolation and permit | Protection setting records, thermographic image history, as-found readings |
| Chillers | Approach temperature, refrigerant, tube condition, compressor health | Removal reduces N+ margin; seasonal capacity constraints | Performance parameter logging, refrigerant register, vendor contract linkage |
| CRAC / CRAH units | Filters, fans, belts, coils, condensate, humidification | High unit count, routine work, but never adjacent units together | Volume PM handling, spatial adjacency rules, filter consumable planning |
| Fire detection & suppression | Detector function, agent pressure, release integrity, interlocks | Statutory; isolation during testing must be time-bounded and recorded | Compliance calendar, certificate storage, isolation start and end times |
| Leak detection | Cable and spot sensor function, false alarm history, zone mapping | Frequently under-maintained because it rarely does anything | Functional test records per zone, alarm history correlation |
| Pumps, valves, pipework | Seal condition, vibration, water treatment, valve exercise | Isolation affects a cooling loop, not a single unit | Loop-level asset hierarchy, water treatment result logging |
Where the task detail matters, the general PM content for these classes is covered elsewhere and applies directly here: electrical preventive maintenance checklists, chiller preventive maintenance, generator, pump and motor PM, and fire and life safety system PM. What changes in a data centre is not the task list, it is the work control wrapped around it.
One asset class deserves a specific note. Leak detection is the system most likely to be quietly broken, because it is passive, it sits under floors and above ceilings, and nothing tells you it has stopped working. A functional test per zone, recorded per zone rather than as one work order for the whole building, is the only way to know it will alarm when it needs to.
6. Criticality classification when almost everything is critical
Standard criticality models struggle in a data centre because the honest answer for most of the asset register is "critical". If everything scores highest, the classification has told you nothing and prioritisation collapses back to whoever shouts loudest. The way out is to score on two axes rather than one.
The first axis is consequence of failure, the usual measure: what happens to the critical load if this asset stops. The second, and the one that makes the model useful here, is redundancy state: how many failures away from load loss is this asset right now. A chilled water pump in an N+1 set has high failure consequence and low immediate exposure while its partner is healthy. The same pump while its partner is out for maintenance has high consequence and high exposure. That is a different asset for the next four hours.
Building that second axis into the CMMS is what turns criticality from a static label into an operational input. In practice it means the system knows the current redundancy state of each group, and the effective priority of a corrective work order rises automatically when the group is degraded. A UPS fan fault is a routine corrective when the parallel module is healthy. It is an emergency when the parallel module is already on bypass. Most CMMS configurations cannot make that distinction, and the good ones can be made to with a calculated priority field driven by a redundancy-state attribute. The underlying scoring method is in asset criticality classification; the data centre addition is the second axis.
The asset hierarchy needs the same treatment. A flat register of UPS units and CRAH units tells you nothing about which path anything sits on. Model the distribution topology: source, path A and path B, board, downstream. It is more work at implementation and it is the difference between a system that can warn you about a conflicting isolation and one that cannot.
7. Maintenance windows and customer SLAs in colocation
In an enterprise-owned data centre the maintenance window is negotiated internally with application owners. In colocation it is negotiated with paying customers who have their own commitments downstream, and that changes the mechanics considerably.
The typical colocation contract obliges the operator to give notice of planned maintenance, often tiered by the risk class of the work, with longer notice for anything that reduces redundancy on a customer's feed. Notice periods of days for routine work and weeks for critical work are common, and customers frequently retain the right to object to a window or request rescheduling. Some large customers require their own change freeze periods to be respected, which in retail-heavy portfolios can remove whole months of the calendar.
The practical consequences for maintenance management are specific.
- Notification is a workflow step, not a courtesy. The CMMS should not allow a critical planned work order to reach released status without a recorded notification event: who was notified, when, by what channel, and what the response was. If notification lives in an operations mailbox, it will eventually be the thing you cannot produce in an audit.
- The window is a constraint the scheduler must respect. Blackout periods, customer freeze windows and permitted maintenance hours belong in the system as calendar constraints, so that a planner cannot innocently schedule switchgear work inside a customer's peak trading freeze.
- Reschedules need a reason code. A PM deferred because the customer objected is a completely different management signal from a PM deferred because the technician was sick or the part did not arrive. If both close as "rescheduled", you lose the ability to see that customer objections are systematically pushing your electrical PM out of compliance.
- PM compliance must be reportable per customer suite. Customers increasingly ask not just whether the site is maintained but whether the plant serving their specific footprint is maintained on schedule. That requires the asset-to-suite relationship to exist in the data model.
Where SLA structure itself needs designing, the general framework in SLA matrix design for FM operations transfers well, with the caveat that data centre availability commitments are usually expressed against power and environmental delivery rather than response and rectification times, so the measurement points differ.
The honest tension nobody solves cleanly
Customer-driven window restrictions and manufacturer-recommended maintenance intervals are in permanent conflict, and there is no configuration that resolves it. Every colocation operator I have discussed this with carries some quantity of overdue or deferred critical PM because the window was not available. The only defensible position is to make the deferral explicit: a recorded deferral with a reason, a risk acceptance by a named person, a revised target date, and a visible count on the management report. What you must not do is let the system quietly reschedule and show green. The deferral is a real risk position and it should look like one.
8. Vendor-delivered maintenance, OEM contracts and holding the warranty
A significant share of critical plant maintenance in a data centre is performed by the OEM or an authorised partner, because the equipment is complex, the warranty requires it, and in some cases only the manufacturer holds the tooling and firmware. That means your maintenance management system is managing a supply chain as much as a workforce.
What I would insist on holding in the CMMS for every vendor-maintained asset:
- The contract record itself, linked to the assets it covers, with scope, frequency, response commitments, expiry and renewal date. Assets falling off contract unnoticed is a common and expensive discovery.
- Scheduled visits as work orders in your system, not only in the vendor's. If the vendor's schedule is the only schedule, you cannot see a missed visit until you look for it.
- The service report as a required closure artefact. The work order does not close until the vendor's report is attached, and someone on your side has reviewed it. Unreviewed vendor reports are where recommendations for remedial work go to die.
- Recommendations raised as follow-on work orders. A vendor report that says a capacitor bank is approaching end of life is only useful if it becomes a tracked item with an owner and a date.
- Firmware and configuration changes recorded as changes. Vendors update firmware during visits and it does not always reach your change register. It should.
- Warranty terms and claim evidence. Serial numbers, commissioning dates, warranty end dates, and the maintenance record that proves the conditions were met.
That last point is worth dwelling on, because warranty is where maintenance record quality converts directly into money. Manufacturer warranties on UPS systems, chillers, generators and switchgear are conditional on documented maintenance at specified intervals by qualified personnel. When a major component fails inside warranty, the claim is assessed against your records. A gap in the PM history, a missing service report, a job closed with no readings, and the claim weakens. The discipline is unglamorous and the payoff is occasionally very large. The mechanics of setting this up properly are in warranty management in asset systems.
The vendor evidence test
Pick any critical asset under OEM contract and try to produce, in under ten minutes and from the CMMS alone: the last three service visits, the reports for each, the recommendations raised and their current status, the firmware revision now running, and the warranty expiry. If you cannot, you do not have a vendor management problem, you have a record-keeping problem that will surface as a warranty rejection or an audit finding.
9. Spares strategy for a facility that cannot wait
Spares management in a critical facility is governed by lead time against consequence rather than by carrying cost. The question is not "how often do we use this part" but "what happens during the weeks it takes to get one".
A practical way to sort the stores:
- Consumables with predictable draw: filters, belts, lamps, treatment chemicals. Standard min/max reorder against actual consumption, which the CMMS should be calculating from issued parts rather than from a number someone set at handover.
- Insurance spares: long lead-time, high-consequence items held not because they will be used but because the facility cannot survive the lead time. Breaker spares, control cards, specific pump assemblies. These need shelf-life monitoring and, for anything with capacitors or batteries, periodic verification that the spare still works.
- Vendor-held stock: parts the OEM commits to holding regionally under contract. Record the commitment and the location, and test it occasionally, because a contractual four-hour part that lives on another continent is a fiction you want to discover before an outage.
- Batteries and shelf-life items: anything that degrades in storage needs a manufactured date, a shelf-life expiry and a rotation rule in the system, or you will eventually install a spare that is already worn out.
The one spares practice I would call non-negotiable in a critical facility is periodic functional verification of insurance spares. A spare control card that has sat in a cupboard for six years and does not work when you finally need it is worse than no spare, because your recovery plan depended on it.
10. Environmental monitoring, evidence, and the audit burden
Environmental conditions in a data hall are monitored continuously by the building systems, against the thermal envelopes published by ASHRAE and whatever tighter internal standard the operator has set. That monitoring is not a CMMS function, and I would not try to make it one. What the CMMS does need is the maintenance-relevant slice of it.
Two things specifically. First, the alarm that becomes work: a condition excursion that requires intervention should generate a corrective work order with the triggering data attached, so the response is tracked and closed like any other job rather than resolved verbally on shift. Second, and less commonly done, the environmental record for the duration of a critical intervention: while a chiller was out for service, what happened to hall temperature. That is the evidence that the work was executed within a safe envelope, and it is exactly what a customer or an investigator asks for afterwards.
The audit burden in this sector is heavier than in almost any other facilities context, and it comes from several directions at once: customer audits and site visits, certification and management-system audits covering information security and service management, insurer inspections, statutory fire and electrical compliance, and internal resilience reviews. Each wants a slightly different cut of the same underlying record.
The pattern that works is to stop treating audits as projects and treat the CMMS output as continuously audit-ready. Concretely that means five things are always true: PM compliance is reportable per asset and per period without manual assembly; every critical intervention has a retrievable approval, procedure revision and completion record; statutory certificates are stored against the asset with expiry tracking; vendor reports and recommendations are in the system with status; and deferrals carry a reason and a named risk acceptance. Sites that hold those five continuously spend hours on an audit. Sites that do not spend weeks, every time, and the preparation work itself distracts the team from operations.
11. DCIM versus CMMS versus BMS: who owns what
This is the question that derails more data centre systems conversations than any other, usually because each vendor category describes its product as the single pane of glass. They are three different systems of record with genuine overlap at the edges, and the useful way to divide them is by the question each one answers.
- BMS answers: what is the plant doing right now, and is it within limits. It is the real-time monitoring and control layer for mechanical and electrical plant.
- DCIM answers: what is installed where, what capacity is consumed and what is left. Space, power, cooling and connectivity capacity, rack and cabinet inventory, customer footprint, and increasingly IT asset lifecycle.
- CMMS answers: what work has been done, what is due, who did it, under what approval, with what result. The maintenance and work management system of record.
| Domain | System of record | Others may consume | Common mistake |
|---|---|---|---|
| Real-time plant status and alarms | BMS | DCIM dashboards, CMMS alarm-to-work-order | Expecting the CMMS to be an alarm console |
| Space, rack and capacity inventory | DCIM | CMMS for asset-to-location linkage | Maintaining a parallel rack register in the CMMS |
| Power chain topology and capacity | DCIM | CMMS for redundancy pairing and isolation logic | Two topologies that disagree with each other |
| M&E asset register and nameplate data | CMMS (with DCIM for IT assets) | DCIM for the capacity view | Splitting one asset across both with no master |
| PM schedules, work orders, history | CMMS | DCIM for a maintenance-status overlay | Using DCIM's light work module and outgrowing it |
| Procedures, approvals, permits, change | CMMS (plus document management) | Service management tool for the change register | Change and work order with no link between them |
| Environmental trend and reporting | BMS or DCIM, depending on the site | CMMS for the slice attached to interventions | Three sources reporting different hall temperatures |
| Spares, stores and purchasing | CMMS or ERP | DCIM for installed-base context | Stores in a spreadsheet beside a capable CMMS |
The overlaps are real and worth naming honestly. DCIM products ship maintenance modules; they are usually adequate for scheduled inspections and inadequate for permit control, vendor contract management, stores and warranty. CMMS products ship condition monitoring integrations; they are not alarm management platforms and should not be asked to be. BMS front ends offer maintenance reminders; those are a convenience, not a maintenance record.
Only three integrations genuinely earn their keep, and I would build them in this order.
- BMS alarm to CMMS work order, filtered hard. Send only the conditions that require a maintenance intervention, with the asset reference, the value and the timestamp. Unfiltered alarm flooding into a CMMS is the single most reliable way to destroy trust in the work queue.
- Runtime and meter data from BMS to CMMS, so meter-based PM triggers on actual hours rather than assumed ones. Generator and pump run hours are the obvious cases.
- Asset and location reconciliation between DCIM and CMMS, with one of them explicitly the master per asset class. Usually DCIM masters racks and IT assets, CMMS masters M&E plant, and location identifiers are shared.
For the controls and building automation side of this picture, which I have deliberately kept out of scope here, see BMS for data centres and critical facilities.
12. What to actually look for when selecting
Very little of the standard CMMS evaluation checklist distinguishes products for this use case. Everyone has mobile, everyone has PM scheduling, everyone has dashboards. The questions that separate a viable critical-facility platform from an unsuitable one are narrower.
- Can it model redundancy relationships and refuse or flag conflicting isolations? This is the single most differentiating requirement and most lightweight products cannot do it without customisation.
- Can a work order carry a mandatory approval chain that varies by change class? Configurable multi-stage approval driven by a work classification, not a single approver field.
- Does it record the procedure revision used, not just attach a file? A text or lookup field that captures document number plus revision at execution time is the minimum.
- Is permit to work native or genuinely integrated? Raised from the work order, blocking closure, with isolation references.
- Can it manage vendor contracts, visits and reports as first-class objects? Not as attachments on a work order.
- Does it handle deferral with reason codes and risk acceptance? And report the deferred position separately from the compliant one.
- Can it report PM compliance sliced by customer footprint? Only relevant in colocation, and hard to retrofit.
- Is the audit trail immutable and exportable? Who changed what, when, with no silent edits.
On products: the enterprise platforms, IBM Maximo, Hexagon EAM, Infor EAM and SAP PM, can all be configured to meet these requirements because they are configurable at the data model level, and the cost is implementation effort and ongoing administration. The mid-market and modern cloud products, MaintainX, Limble, Fiix, UpKeep and eMaint, are considerably easier to deploy and are strong on mobile execution and adoption, but tend to be weaker on multi-stage conditional approval, redundancy modelling and vendor contract depth. Planon is strong where the data centre sits inside a broader corporate real estate portfolio. There are also purpose-built critical facility management products aimed specifically at this sector; they fit the process well and you should examine their integration story and their asset data model carefully before assuming they scale with the estate.
The selection error I see most often is choosing on user experience alone. Adoption matters and a product technicians refuse to use is worthless, so ease of use is a real requirement. But in this sector it is a second-order requirement. A platform with an excellent mobile app that cannot enforce a conditional approval chain or model a redundancy pair will pass the demo and fail the first audit.
The idea to walk away with
In a data centre, the maintenance system earns its place by controlling the risk that maintenance itself introduces. That reframing changes every design decision. You model redundancy so the system can tell you when an isolation is unsafe. You classify change so approval effort matches consequence. You hold procedure revisions against assets so nobody executes a stale script. You capture evidence continuously because it will be demanded. You manage vendors as a supply chain because they perform much of the work and hold the warranty. And you keep DCIM, BMS and CMMS in their own lanes, integrated at the three points that matter, instead of buying whichever one claims to replace the other two.
None of that is exotic technology. It is data model discipline and process design, applied with the assumption that the next outage will be caused by a well-intentioned person doing planned work on the wrong side of a redundant pair.
Final thoughts
If you are standing up or rebuilding maintenance management for a critical facility, the sequence I would recommend is unglamorous and reliable. Get the asset hierarchy and the distribution topology right first, including redundancy pairing, because almost every capability above depends on it. Then build the change classification and approval workflow, because that is the control that prevents the outages. Then wire permit to work and procedure revision capture into the work order. Then vendor contracts, visits and warranty. Then the integrations, filtered hard. Reporting comes last and comes easily if the four layers beneath it are sound.
Resist the pull to start with dashboards, because dashboards are what gets demonstrated and asked for. A beautiful availability dashboard sitting on top of an asset register that does not know which UPS backs up which is a risk, not a control. The order matters more than the product.
Disclosure
Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.
Building maintenance management for a critical facility?
Independent advisory on asset hierarchy and redundancy modelling, change control design, procedure and permit integration, vendor contract structure, and the DCIM/CMMS/BMS boundary. 22+ years across enterprise CMMS, EAM and CAFM implementations. No reseller arrangements.
Book a conversationRelated reading: BMS for data centres and critical facilities, CMMS for facilities management, Asset criticality classification, Permit to work integration, Warranty management, PM examples across industries.
Muhammad Abbas
CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.
Work with me