mail@mabbaz.com Abu Dhabi, UAE

CMMS · Manufacturing · Plant Systems

CMMS for Manufacturing Plants

Most CMMS products are designed around a building, where maintenance owns the asset and books its own access. A production plant inverts that: operations owns the asset, maintenance requests access, and the system has to support a negotiation rather than assume a calendar. This is a practitioner's guide to configuring and integrating a CMMS for a factory: line-shaped asset hierarchy, downtime reason codes that feed OEE, shutdown work modes, MES and historian integration, shift handover, and the requirements plant managers ask for that generic systems quietly fail.

Muhammad Abbas September 25, 2026 ~23 min read

There is a specific moment in a manufacturing CMMS implementation where the project either turns the corner or quietly begins to fail. It is the moment somebody from production sits in a design workshop, looks at the proposed asset hierarchy, and says: that is not how the plant works. The hierarchy on the screen is organised by building, floor and room, because that is how the software's demonstration data was organised and nobody questioned it. The plant is organised by line, cell and station, because that is how material flows and how output is counted. Those two shapes cannot be reconciled after go-live, and almost every problem a manufacturing site later has with its CMMS traces back to a handful of decisions of exactly that kind. This guide is about those decisions.

The message up front: a manufacturing CMMS is not a facilities CMMS with different asset names. It has to model the production line as the primary hierarchy, capture downtime against a controlled reason-code structure that reconciles with the production system's numbers, treat shutdowns as a distinct work mode rather than a busy week, and pull runtime and fault data from the control layer instead of asking technicians to transcribe hour meters. Get those four right and an ordinary mid-market product will serve a plant well. Get them wrong and no amount of licence spend will fix it.

This article is about the software: what the system must support, how it should be configured and what it has to integrate with. The maintenance programme itself, the criticality logic, the equipment-class PM content, the shutdown scope discipline and the schedule structure, is a separate subject and I have written it up in full in preventive maintenance for industrial and plant equipment. If you have not settled the programme, settle that first. Configuring a system around an undefined programme is how you end up encoding guesswork.

1. The production tension the system has to model

Start with the constraint that shapes everything else. In a facility, maintenance largely controls when work happens. In a plant, it does not. Access to a machine is granted by whoever owns the production plan, and the CMMS is either a tool that supports that conversation or an irrelevance that generates work orders nobody can execute.

Most out-of-the-box configurations assume maintenance control. A PM generates on its due date, lands in a queue, and is expected to be done. When it is not done, it goes overdue, and the compliance report turns red. Six months in, the planner is manually deferring hundreds of work orders a month and has stopped believing the numbers. The system is not wrong about the dates; it was never told that access is conditional.

What a plant configuration needs instead:

  • A "running or shutdown" attribute on every PM task. This is the single highest-value custom field in a manufacturing CMMS. Tasks that can be executed with the line running go into the weekly schedule. Tasks that need isolation, entry or stoppage go into a shutdown pool and are never scheduled against a calendar date at all. Without this split the compliance figure mixes two populations with completely different achievability and tells you nothing.
  • A production-access request that is a real record. When maintenance needs twenty minutes on a machine, that request should exist as an object in the system with a requested window, a requesting planner, a production approver and an outcome. When it is refused, the refusal is recorded against the asset. That record is what turns "production never gives us time" from a complaint into a report you can take to a plant meeting.
  • Deferral with a reason code, not silent overdue. A PM that could not be executed because the line was running is a different event from one nobody got to. Configure a deferral reason list (production access refused, parts not available, labour not available, superseded by corrective work, deferred to shutdown) and make it mandatory. The distribution of those reasons over a quarter is the most useful diagnostic a plant maintenance manager can hold.
  • A grace window per criticality band, agreed with production. A monthly task with a plus or minus seven day window is honest and achievable. A monthly task due on the fourteenth or it is late is a fiction in a plant. Agree the windows once, configure them, and then hold the compliance line hard inside them.
The configuration test

Ask the vendor to show you, in their system, how a planner requests a two-hour window on a specific machine, how production approves or refuses it, and where the refusal appears in a report. Most demonstrations cannot do this without a workaround, because the product assumes the window is yours to take. How gracefully they handle the question tells you how much of your plant reality the product has actually met.

2. Asset hierarchy shaped by the production line

The hierarchy is the decision with the longest tail. Everything you will ever want to report, roll up or filter depends on it, and restructuring it after two years of work-order history is painful enough that most sites never do. The general design principles are in asset hierarchy design for CAFM and EAM; what follows is the manufacturing-specific shape.

The primary hierarchy should follow the production flow, not the geography:

Site
  ↓
Production area / value stream  (Mixing, Forming, Packing)
  ↓
Line  (Line 2)  ← the OEE reporting level
  ↓
Cell / station  (Line 2 filler, Line 2 capper)
  ↓
Equipment / functional location  (capper drive unit)
  ↓
Component / maintainable item  (gearbox, servo, sensor)

Two specifics matter more in a plant than anywhere else. First, the line must be a real node, not a text field on the asset. Almost every question a plant manager asks is asked at line level: downtime by line, cost by line, PM compliance by line. If the line is an attribute rather than a parent, every one of those reports becomes a filter that somebody will eventually get wrong.

Second, encode position in the flow and constraint status as asset attributes. At minimum: sequence position within the line, whether the asset is the current constraint, whether there is buffer downstream, and whether a bypass or parallel route exists. These are the fields that let you answer the only question that matters during a breakdown, which is whether this stoppage stops the line, and they are the fields that make a criticality ranking mean something. The scoring method behind them sits in asset criticality classification. The configuration point is that the constraint flag must be editable and reviewed, because the bottleneck moves when the product mix moves, and a hard-coded criticality band will be silently wrong within a year.

Keep geography as a secondary classification rather than the spine. Building, floor and room still matter for access, contractor induction and utilities, so hold them as attributes or a parallel location dimension. In IBM Maximo and SAP PM this is the classic functional-location-plus-classification split; in mid-market products such as Fiix, Limble, eMaint or MaintainX you generally get one hierarchy plus tags, in which case the hierarchy goes to the production flow and geography becomes tags. Do not spend the hierarchy on the building.

3. The manufacturing capability requirements table

When a plant evaluates CMMS products, the generic feature checklist is close to useless because every product ticks it. The requirements that actually separate a system that will work in a factory from one that will not are narrower and more awkward to demonstrate. This is the list I would put in front of a vendor, with the reason each one exists. The general module taxonomy, for context, is in maintenance management system core modules explained.

Capability What it must actually do Why a plant needs it Priority
Line-based hierarchy Multi-level parent/child with the line as a rollup node, and cost and downtime aggregating up it Every plant report is asked at line level Must
Running / shutdown task flag Attribute on the PM task that routes it to a weekly schedule or a shutdown pool Separates achievable work from work needing an outage Must
Downtime capture on the work order Start and end of equipment stoppage as distinct fields from labour time, plus a mandatory reason code Labour hours are not downtime, and conflating them destroys the Pareto Must
Controlled reason-code hierarchy Two or three level picklist, mandatory, no free text substitute, editable only by an owner The only way downtime analysis stays comparable over time Must
Meter-driven PM from an external source Accept running hours, cycles or units via API or file, and trigger PM on the imported value Manual meter transcription fails above roughly fifty assets Must
Shutdown / project work mode A parent event with scope list, scope freeze, readiness state per job, and durations that can be sequenced Shutdowns are a distinct planning mode, not a busy week Must
Shift-aware work management Shift calendar, shift stamped on every transaction, handover view of open and in-progress work Continuous operations hand work across crews mid-job Must
Failure coding to Problem / Cause / Action Three linked fields, valid-value lists constrained by equipment class Without it you cannot move from downtime to root cause Must
Spares with lead time and criticality Lead time as a maintained field, reorder logic that uses it, insurance-spare flag, part-to-asset link Long-lead production-critical parts are the real stockout risk Must
Audit trail and record integrity Immutable transaction log, attachment retention, no silent overwrite of completion data Regulated manufacturing evidence, and PM history you can trust Must in regulated sites
Electronic signature and record controls Signature meaning, signer identity, timestamp, reason for change on any amendment Pharma, medical device and some food operations require it Must in regulated sites
Permit and isolation linkage Work order cannot progress to execution without the linked permit in the right state Outage work with contractors on site is the highest-risk window Should
Contractor access and competency Contractor as a resource type, induction and certificate expiry blocking assignment Turnaround labour is mostly not yours Should
Calibration management Instrument register, interval, as-found and as-left values, certificate attachment, out-of-tolerance workflow Quality loss and audit exposure concentrate here Should
Mobile execution offline Full work order execution with no signal, queued sync, no data loss on conflict Plant floors, basements and steel structures kill wifi Should
Open integration surface Documented REST API covering assets, work orders, meters and parts, plus webhooks or an event queue MES, historian and ERP all need to talk to it Must
Reporting at line and shift grain Downtime and cost pivoted by line, cause, shift and period without exporting to Excel This is the report the plant manager reads Must

A note on how to use that table. Do not send it as a questionnaire, because every vendor will answer yes to every row. Pick the six rows that matter most on your site and ask for each to be demonstrated live, in a configured system, with your terminology. The gap between "supported" and "supported without a consultant writing a customisation" is where implementation budgets disappear.

4. Downtime reason codes: the structure that makes or breaks the analysis

If I had to name the single configuration decision that determines whether a manufacturing CMMS produces useful analysis or merely stores records, it is the downtime reason-code structure. Get it right and in twelve months you have a Pareto that tells the plant exactly where to spend its reliability effort. Get it wrong and you have a chart where forty percent of downtime sits under "other" or "mechanical fault", which is worse than no chart because people act on it.

The structure that works is three levels, each answering a different question, with the third level constrained by the second. Level 1 is who owns the loss, which is the level production and finance care about. Level 2 is the loss category. Level 3 is the specific cause, which is the level maintenance acts on.

L1 Ownership L2 Category L3 Specific cause (examples) Counts against OEE as
Equipment
maintenance owns
Mechanical failure Bearing · gearbox · belt or chain · seal or leak · structural · wear part Availability (breakdown)
Electrical failure Drive or inverter · motor · supply or protection · termination · panel cooling Availability (breakdown)
Control and instrument Sensor · PLC or IO · network · HMI · calibration drift · software fault Availability (breakdown)
Minor stop Jam · misfeed · sensor false trip · reject ejection · guard trip Performance (small stops)
Production
operations owns
Changeover Product change · format change · clean down · first-off approval wait Availability (setup)
Operating Operator absent · speed reduced deliberately · process adjustment · startup ramp Performance
No demand No orders · planned idle · break or shift gap Excluded (not planned time)
Supply
supply chain owns
Material Raw material short · packaging short · out-of-spec material · upstream starve Availability (idling)
Downstream Blocked by downstream · warehouse full · no pallets or transport Availability (blocked)
Utility
shared
Services Compressed air · steam · chilled water · power · extraction or HVAC Availability (breakdown)
Quality
quality owns
Quality hold Out-of-spec output · deviation investigation · rework · QA release wait Quality, and Availability if stopped
Planned maintenance
maintenance owns
Scheduled work Planned PM · agreed access window · shutdown · statutory inspection Availability, tracked separately from breakdown
Rules that keep this usable: L3 is mandatory, L1 and L2 derive from it automatically · no "other" at L3 on any code that occurs more than twice a month, split it instead · the list is owned by one named person and changes go through a monthly review, never by a supervisor adding a code at two in the morning · a code is retired, not deleted, so history stays readable · total codes stay under roughly eighty, because a list a shift leader has to scroll past eighty entries will be answered with whichever code is nearest the top.

Three configuration details that decide whether this holds up in practice. First, capture the reason at the stoppage, not at work-order closure. A code applied by a planner three days later from a scribbled note is a guess. It should be selected by the operator or shift leader on the line terminal or the mobile device, within the shift.

Second, separate downtime duration from labour duration. They are different numbers and conflating them is the most common error I see. Equipment was down four hours; a technician spent ninety minutes on it and two hours waiting for a part. Both facts matter and each answers a different question. If the system only has a labour-hours field, the downtime analysis is already broken.

Third, reconcile monthly against production's own downtime total. Pick one line, one month, and compare the CMMS downtime sum against the production log. If they disagree by more than a few percent, find out why before anybody builds a Pareto. On the failure coding itself, which is a related but distinct structure, the Problem, Cause, Action approach is covered in the work-order discipline described in work order types in a CMMS.

Where the reason-code idea breaks down

A reason-code structure only produces trustworthy analysis if the people applying it believe the analysis will be used fairly. The moment L1 ownership is used to allocate blame in a plant meeting, coding behaviour changes: production codes to Equipment, maintenance codes to Production, and within two months the data is political rather than factual. I have seen well-designed code structures become useless for exactly this reason. If your plant culture cannot hold an ownership dimension neutrally, collapse L1 and run the analysis on L2 and L3 only. Slightly less insight, honestly collected, beats a richer structure everyone games.

5. What the CMMS owes OEE (and what it should not try to own)

OEE is calculated in the production system, and the CMMS should not try to become the OEE tool. Plenty of implementations attempt it, usually because the MES project was postponed, and the result is a maintenance system holding production data that production does not trust and does not use. The right division of labour is narrower and much more achievable.

What the CMMS owns:

  • Equipment-attributable downtime, by asset, cause and shift. The availability loss that belongs to maintenance, coded properly, timed properly, linked to a work order and a failure code. This is the CMMS contribution to the OEE conversation.
  • Minor stops and chronic loss, where they are captured at all. The most valuable and most neglected input. Micro-stops rarely generate work orders, so they are invisible to maintenance while quietly costing more output than headline breakdowns. Configure a lightweight path for logging them, ideally a one-tap entry on a line terminal that aggregates rather than raising a work order each time.
  • Planned maintenance downtime, tracked separately. Planned and unplanned stoppage must never sit in the same bucket, because the whole argument for PM depends on showing one falling as the other is spent.
  • The maintenance causes hiding inside quality and performance loss. When the CMMS holds calibration records and instrument history, and quality holds scrap by cause, the join between them exposes drift-driven scrap that neither system sees alone.

What the CMMS should leave alone: units produced, ideal cycle time, quality yield, order and schedule data. Consume those, do not master them. The integration should be one-directional for those fields, with the MES or the line system as the source of truth, and the CMMS storing only what it needs to contextualise a work order.

The one thing worth insisting on is a shared definition of planned production time. Maintenance and production frequently disagree about whether a planned changeover, a break, or an idle shift counts inside the denominator, and when they disagree, their availability numbers differ and every subsequent conversation is an argument about arithmetic. Agree the definition in writing during design, document it in the configuration, and revisit it only deliberately. The KPI framework covers the broader discipline of defining a measure before automating it, which applies here more sharply than anywhere.

6. Shutdowns and turnarounds as a distinct work mode

A shutdown is not a week with a lot of work orders in it. It is a different planning mode with its own scope object, its own approval gates, its own resource model and its own critical path, and a CMMS that cannot represent that will push the entire turnaround into Microsoft Project or Primavera, at which point the shutdown history never reaches the maintenance record.

What the system needs to support:

  • A shutdown event as a parent record. Named, dated, with a scope list of child work orders drawn from PM falling due, backlog, condition findings, statutory inspections and projects. Every job in the shutdown is reachable from one place.
  • Scope freeze as a system state, not an email. After the freeze date the scope list locks, and additions require a named approver and a recorded justification. If the system cannot enforce this, scope creep is unmanaged by construction.
  • Readiness state per job. Parts staged, permit drafted, isolation identified, labour named, method agreed. Five flags, rolled up into a readiness percentage for the whole shutdown. This single view is the most valuable screen in a turnaround and most CMMS products do not have it out of the box, so plan to build it as a custom status set or a report.
  • Durations, dependencies and resource loading. Enough scheduling capability to see whether you have promised four hundred hours of work to a crew that can deliver two hundred and forty. If the product cannot do dependencies, accept the export to a scheduling tool, but insist that completion and findings come back into the CMMS so the asset history is intact.
  • Long-lead procurement linked to the scope. A shutdown job whose part has a fourteen-week lead needs to be visible as a risk sixteen weeks out, not discovered at the freeze. That means the scope list has to see requisition and expected-delivery status, which usually means the CMMS to ERP integration has to carry purchase-order status back.
  • A post-shutdown close-out. Planned versus actual duration per job, work not completed and carried, findings raised, and the scope that was added after the freeze. Without this the next shutdown repeats the same estimating errors.

7. Integration with MES, SCADA and the historian

This is where a manufacturing CMMS differs most from a facilities one, and it is the integration work most sites underestimate. There is real data in the control layer that the maintenance system needs, and there is far more data that it should be kept away from.

The four flows that genuinely earn their place:

1. Runtime meters  historian → CMMS  (hours, cycles, units; daily or shiftly)
    drives meter-based PM without transcription

2. Stoppage events  MES / line system → CMMS  (start, end, asset, initial reason)
    seeds the downtime record; human confirms the code

3. Fault and alarm codes  SCADA / PLC → CMMS  (filtered, not raw)
    attaches machine-side diagnostics to the work order

4. Condition exceptions  monitoring → CMMS  (threshold breach, trend alarm)
    becomes a work request, not an email

    CMMS → MES: asset status, planned outage windows only

The architectural point that saves projects: integrate at the historian, not at the PLC. A CMMS should never hold a direct connection into a control network. Read from the time-series historian or a dedicated integration layer, on a scheduled pull, across a firewall boundary, with the CMMS on the enterprise side. The mechanics of that boundary, the tag mapping and the aggregation are covered in SCADA and historian integration and the work-order side in SCADA to CMMS integration.

Some specific decisions worth making deliberately:

  • Aggregate meters, do not stream them. A CMMS needs a running-hours total per asset per day. It does not need a reading every five seconds, and loading high-frequency data into a maintenance database degrades it for no benefit. Aggregate in the historian, push a daily or per-shift delta.
  • Filter alarms hard before they reach the CMMS. An unfiltered alarm feed will generate thousands of work requests a week and the queue will be abandoned within a fortnight. Only alarms that have been assessed as actionable, on assets that matter, with de-duplication and a minimum persistence, should cross the boundary. Start with ten alarm types, not ten thousand.
  • Never auto-close a work order from a control signal. A machine restarting is not evidence the fault was fixed. Machines restart because somebody reset a trip. Auto-creation with human confirmation is sound; auto-closure destroys history.
  • Map tags to assets in a maintained table, not in code. Tag names change when a machine is modified. If the mapping lives inside an integration script, every modification becomes a development request. Hold it as configuration that an engineer can edit.
  • Decide what happens when the link is down. Meters queue and catch up, ideally. Stoppage events buffer. If the design silently drops data during an outage, meter-based PM will drift and nobody will notice for months.

If the plant is also pursuing condition monitoring, the same boundary carries it, and the case for where that investment actually pays is set out in predictive maintenance in manufacturing. The integration principle is identical: an insight that arrives anywhere other than the technician's work queue will be ignored.

Sequence the integrations

Build runtime meters first. It is the smallest, most reliable flow, it delivers immediate value by removing manual transcription, and it proves the boundary works. Stoppage events second. Alarms and condition exceptions last, and only after somebody has done the work of deciding which ones are actionable. Sites that start with the alarm feed because it looks impressive in a demonstration almost always end up switching it off.

8. Spares configuration for long-lead, production-critical items

The general MRO inventory discipline is covered in spare parts and MRO inventory in a CMMS. The manufacturing-specific configuration problem is narrower: standard reorder logic is built around consumption rate, and the parts that hurt a plant most are the ones with almost no consumption and a lead time measured in months.

A min/max model driven by usage history will set the reorder point on a large gearbox to zero, because it has never been consumed. That is arithmetically correct and operationally disastrous. The fields and rules that fix it:

  • Lead time as a first-class maintained field, with a review date. Not a note in the description. Refresh it annually with the supplier, because quoted lead times have moved substantially in recent years and a number entered at handover is fiction.
  • An insurance-spare flag that exempts the part from consumption-based reorder logic and instead drives an annual holding review. These parts are held because the plant cannot survive the lead time, not because they will be used. Each one needs a written justification linked to the asset criticality, stored on the part record so the next cost-cutting review can see why it exists.
  • Part-to-asset linkage that reaches the constraint flag. You need to be able to ask: which parts, if unavailable, would extend a stoppage on the bottleneck. That query is impossible unless parts are linked to assets and assets carry a constraint attribute, which is why the hierarchy decision in section two reaches this far.
  • An obsolescence register. A flag and an end-of-life date on drives, controllers and instruments whose manufacturer has issued a notice. Lead time on an obsolete part is effectively infinite, and these should surface on a quarterly report rather than being discovered during a breakdown.
  • Kitting against the work order. Parts reserved and staged before the window opens, not fetched during it. In a plant where the window is two hours and was hard to get, a technician walking to stores is a material loss.
  • Repairable and rotable handling. Large rotating equipment often runs on a repair-exchange basis. If the system cannot track a component out for refurbishment as an asset in a repair state, that population will be managed on a spreadsheet and periodically lost.

9. Shift handover and multi-crew work

A facilities CMMS quietly assumes a day shift. Plants run two, three or continuous shifts, work crosses crew boundaries mid-job, and the handover is where information is lost. The configuration requirements are unglamorous and consistently overlooked at design time.

  • A shift calendar the system actually knows about, so that response and completion clocks respect shift patterns rather than counting a night shift as elapsed idle time, and so scheduling does not assign work to a crew that is not on site.
  • Shift stamped on every transaction. Not derived from a timestamp at report time, which breaks on night shifts spanning midnight and on daylight-saving changes. Stored on the record. Without it, downtime by shift, the report plant managers ask for most consistently after downtime by cause, cannot be produced reliably.
  • A handover view, not a handover document. Open work on my area, current status, what was attempted, what is waiting, what is isolated and must not be re-energised. If the handover lives in a notebook or a WhatsApp group, the CMMS is not the system of record for work in progress and its history will have holes.
  • Progress notes as timestamped, attributed, append-only entries. A single free-text field that each shift overwrites is the most common way plants lose diagnostic history. The third shift needs to read what the first two tried.
  • Isolation state visible and persistent across shifts. A machine left isolated at the end of a shift is a safety matter, and the isolation record must be prominent in the incoming crew's view rather than buried in a permit attachment. The linkage pattern is described in permit to work integration with a CMMS.

10. Contractor management during outages

During a normal week a plant maintenance team is mostly its own people. During a turnaround, the majority of the labour on site may be contractors, many of them unfamiliar with the plant, working in parallel in congested areas, under permits, on a compressed timetable. That is both the highest-risk and the least-well-supported scenario in most CMMS configurations.

What the system needs to handle:

  • Contractor personnel as assignable resources with expiring attributes. Induction date, competency certificates, permit qualifications, insurance validity. The system should refuse to assign work to a person whose induction or certificate has lapsed. This is the control that actually works, because it operates at the point of assignment rather than as a reminder nobody reads.
  • Scope and progress visible per contractor package. During a shutdown you need to know which jobs a given contractor holds, their readiness state and their progress, without assembling it from twelve work orders.
  • Permit linkage that blocks execution. Not a reminder, a hard gate: the work order cannot be progressed to in-progress without its permit in an issued state. Contractor work in an outage is exactly where soft controls fail.
  • Findings and completion captured by the contractor, in the system. If contractor work is closed out on paper and typed in afterwards, the detail is lost and the asset history for the shutdown is thinner than for any ordinary week, which is the opposite of what it should be. Give them limited mobile access with a constrained view.
  • Measurement against the contract. Planned versus actual hours and duration per package, defects found on handback, and work returned. Without it, contractor performance is judged by impression and the same underperforming package holder is engaged for the next turnaround.

11. Compliance and traceability in regulated manufacturing

In pharmaceutical, medical device and much of food manufacturing, the CMMS moves from being an operational tool to being a system that produces regulatory evidence, and that changes both the requirements and the cost of the implementation considerably. If you are in one of those sectors, this section is not optional reading and it is not a bolt-on phase.

The requirements that drive real configuration and cost:

  • Validation of the system itself. A CMMS used to demonstrate that qualified equipment is maintained and calibrated is generally a validated system: specification, qualification, traceability matrix, controlled change management afterwards. That last part is the one people underestimate. In a validated environment every configuration change carries a documented assessment, so the cheerful reconfiguration culture that works elsewhere is not available to you. Design more carefully up front, because iterating later is expensive.
  • Electronic records and signatures. Signature meaning, signer identity, timestamp, and a reason recorded for any amendment to a completed record. Records must not be silently overwritable, and the audit trail must be independent of the user, including administrators.
  • Calibration with as-found and as-left values. Not a pass or fail tick. The measured values, the standard used and its own traceability, plus a defined out-of-tolerance workflow that triggers an impact assessment on product made since the last good calibration. The quality-management foundations for measurement traceability sit with the standards bodies; the catalogues are at ISO for quality and measurement management and IEC for electrical equipment and safety.
  • Deviation and change control linkage. A maintenance event on qualified equipment may require a deviation record or a change control in the quality system. The CMMS does not need to own those, but it needs to reference them, so an auditor can walk from the work order to the quality record without a human bridging the gap.
  • Batch and product traceability on the maintenance side. If an instrument is found out of tolerance, you need to know what was produced on that line since it was last verified. That requires the maintenance record and the production record to share an asset identity and a reliable time base, which is one more reason to get the hierarchy and the shift stamping right.
  • Training and competency records tied to task assignment. In regulated sites, who performed the work and whether they were qualified to perform it is part of the evidence, not an HR matter.

A practical observation on food manufacturing specifically, since it sits in an awkward middle ground: the formal validation burden is usually lighter than pharma, but hygiene-driven requirements bite harder in ways generic systems handle badly. Food-grade lubricant control as a hard constraint on the part record, foreign-body risk requiring tool and part accountability on the work order, and clean-down as a distinct work type with its own verification. These are configuration items, not training slides, and if the system cannot enforce a food-grade lubricant constraint at issue then it will be enforced by memory.

The honest cost of regulated implementation

A validated CMMS implementation in a regulated plant typically costs two to three times the unvalidated equivalent in effort, and the ongoing change overhead persists for the life of the system. I am not going to pretend otherwise, and I would be cautious of anyone who quotes a regulated implementation at the same rate as a general one. The corollary is that scope discipline matters much more: every field you add is a field you will validate and re-validate. In regulated sites the leanest configuration that meets the requirement is not a compromise, it is the correct engineering answer.

12. The reports plant managers actually ask for

Maintenance teams tend to build reports about maintenance. Plant managers want reports about production loss. The difference is worth understanding, because a CMMS that only produces the first set will be regarded as a maintenance department tool rather than a plant tool, and its budget conversations will go accordingly.

The reports I would configure in the first month, in the order plant managers ask for them:

  • Downtime hours by cause, Pareto, rolling three months. L3 reason codes, ranked, with the top five carrying an owner and an action. The rolling window matters: a single month is noise and a year hides recent change.
  • Downtime by line, with planned and unplanned separated. Ranked by line so attention goes where the loss is, and separated so a line with heavy planned work does not look like a reliability problem.
  • Downtime by shift. Uncomfortable and consistently revealing. Persistent differences between shifts on the same equipment are almost never equipment problems; they are practice, skill or handover problems, and the report is how you find them.
  • Unplanned downtime on the constraint asset, trended. The single number the plant should manage, because an hour lost at the bottleneck is an hour lost by the plant.
  • PM compliance split by criticality band and by running versus shutdown. Four numbers instead of one. Ninety percent overall hiding fifty-five percent on critical shutdown tasks is a failing programme wearing a good headline.
  • Repeat failures: same asset, same failure code, within ninety days. A short list, reviewed monthly. This is the cheapest reliability report there is and very few plants run it.
  • Deferral reasons for the period. Why planned work did not happen, distributed by cause. This is the report that turns an access argument into a decision.
  • Backlog in crew-weeks, by trade, with age. Whether the resourcing matches the workload, in a unit a plant manager understands.

One structural recommendation: put a maintenance measure inside the production review rather than holding a separate maintenance review. One scoreboard, read by both functions, is worth more than two accurate ones read separately. It is also the quickest route to the access negotiation in section one becoming a planning conversation instead of a standoff.

13. What actually goes wrong (the honest section)

Across manufacturing CMMS implementations the failure patterns are depressingly consistent, and three of them account for most of the disappointment.

Maintenance and production run separate systems of record. Production logs downtime in the MES or on a line spreadsheet. Maintenance logs work orders in the CMMS. The two numbers never agree, and because they never agree, neither side trusts the other's report, so every conversation about reliability becomes a conversation about whose data is right. This is the most common structural failure in the category and it is not a technology problem. It is that nobody was made accountable for one reconciled downtime figure. The fix is unglamorous: agree a single source for the stoppage event, usually the production system because that is where the operator already is, integrate it into the CMMS to seed the work order, and reconcile the totals monthly with both parties in the room. Until the two systems produce the same number, no analysis built on either is worth much.

Downtime reasons are captured inconsistently, so the Pareto is meaningless. The structure from section four is designed well, and then a shift leader at two in the morning with a line down and a queue of product cannot find the right code, so they pick the nearest plausible one. Multiply that by three shifts and eighteen months and the top of your Pareto is an artefact of picklist ordering. The specific failure modes I see repeatedly: too many codes, so nobody scrolls; a free-text field alongside the picklist, so the real information goes into the text and never into the analysis; codes applied at closure by a planner rather than at the event by the person who saw it; and no owner, so codes accumulate until there are three variants of the same cause. Every one of those is a configuration and governance decision, made at design time, by someone who was optimising for completeness rather than for a tired shift leader with one hand on a machine.

PM compliance is sacrificed to the production schedule, every single month. This is the one nobody writes on a slide. The PM programme is designed, loaded and generating, and then month after month the critical tasks needing an outage are deferred because the line is running and the order book is full. Compliance on running tasks stays respectable, compliance on shutdown tasks quietly decays, and eighteen months later there is a significant failure on an asset whose internal inspection has been deferred eleven times. Nobody decided to accept that risk; it accumulated one reasonable deferral at a time.

The system cannot solve that, but it can stop it being invisible, and that is the realistic ambition. Make the deferral reason mandatory. Count deferrals per task and surface any task deferred more than twice consecutively as an exception with a named owner. Report shutdown-task compliance separately so it cannot hide inside the headline. And escalate a third consecutive deferral on an A-band task to a decision that somebody signs, rather than a status that somebody changes. That converts risk accumulation into a series of explicit, attributable choices, which is the most a maintenance system can honestly do about an organisational pressure it did not create.

A fourth pattern worth naming briefly, because it is so common: the plant buys a CMMS to fix a programme that was never designed. The system then generates the same wrong frequencies more reliably, makes the unnecessary tasks harder to delete, and renders the compliance gap visible to an audience that did not previously see it. That last effect is genuinely useful, but it is not what was bought. If you are at the selection stage and the programme underneath is not settled, the buyer's introduction to CMMS sets out what the software does and does not do before you commit to a shortlist.

The idea to walk away with

A CMMS in a manufacturing plant has one job that a facilities CMMS does not: it has to work inside a negotiation it does not control. Maintenance does not own access to the equipment, so the system has to make the negotiation visible. That means splitting every task by whether it needs an outage, recording access requests and refusals as data, making deferrals explicit and attributable, and reporting shutdown compliance separately from running compliance.

Everything else follows from four configuration decisions taken before go-live: a hierarchy shaped by the production line with the line as a real node and the constraint flagged as an attribute; a three-level downtime reason-code structure that is mandatory, short, owned and applied at the event; shutdowns modelled as a distinct work mode with a scope object, a freeze state and a readiness view; and runtime meters and stoppage events flowing in from the historian and the MES rather than being transcribed by hand. Those four are worth more than any feature comparison, and all four are decisions your implementation team makes, not decisions the vendor makes for you.

Final thoughts

The manufacturing sites with the most useful maintenance systems I have seen are rarely the ones running the most capable product. They are the ones where the downtime number in the CMMS matches the downtime number in the production report, where the reason-code list is short enough that a tired shift leader picks the right one, where the planner knows the production schedule, and where a task deferred three times becomes a decision with a name on it rather than a status nobody reads.

If you are starting this work, I would resist the urge to configure broadly. Take one line, ideally the one with the constraint on it. Build the hierarchy properly for that line, down to component level on the critical assets. Configure the reason codes and use them for a month on that line alone. Integrate runtime meters for its assets. Reconcile the downtime figure with production. When those five things hold on one line, the pattern is proven and rolling it across the plant is largely repetition. Configure the whole site at once and you will discover the hierarchy problem in month nine, with two thousand work orders already hanging off it.

Disclosure

Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.

Selecting or configuring a CMMS for a production plant?

Independent advisory on manufacturing CMMS requirements, line-based hierarchy design, downtime reason-code structures, MES and historian integration, and shutdown work management. 22+ years across manufacturing, utilities, oil and gas, government and facility operations. No vendor margins, no reseller arrangements.

Book a conversation

Related reading: Preventive maintenance for industrial and plant equipment, What is a CMMS: a complete buyer's introduction, Maintenance management systems: core modules explained, SCADA to CMMS integration, Predictive maintenance in manufacturing, Asset hierarchy design for CAFM and EAM.

Muhammad Abbas

CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.

Work with me
MAbbaz.com
© MAbbaz.com