mail@mabbaz.com Abu Dhabi, UAE

BMS · Controls · Lifecycle

BMS Maintenance, Servicing and Lifecycle

Everybody maintains the chillers. Almost nobody maintains the system that controls them. A building management system is itself an asset, with sensors that drift, controllers that age out of support, programs that only exist on one engineer's laptop, and a service contract whose exclusions you probably have not read. This is a practitioner's guide to BMS maintenance, service contract structures, obsolescence and the upgrade paths that follow.

Muhammad Abbas September 25, 2026 ~22 min read

Walk into a plant room on a site that has run for eight or ten years and look at the BMS front end. Count the manual overrides. Count the points in alarm that nobody has acknowledged since a handover two contractors ago. Count the zones running twenty four hours because somebody forced them one summer and never put them back. That is what an unmaintained building management system looks like, and it rarely looks broken: it looks like a working system that nobody trusts, which is a quieter and far more expensive failure. BMS maintenance is the most consistently neglected programme on a site for a simple reason. The BMS almost never stops. Mechanical plant announces its neglect; controls just drift.

The message up front: treat the BMS as an asset on your register with its own PM tasks, its own criticality, its own spares and its own end-of-life date. Then settle three questions before you sign anything: who holds the engineering licences, who holds the source graphics and controller programs, and how long the installed controller generation will remain supported. Those three answers determine whether you have a maintainable system or a permanent dependency.

1. The BMS is an asset, and it is almost never on the register

Open the asset register on most sites and you will find chillers, AHUs, pumps, fans, generators, lifts, fire panels. You will often find not a single record for the building management system. Controllers, field devices, supervisory server, operator workstation, the switches carrying controls traffic: all capital plant with a service life, a support horizon and a replacement cost, and none of it planned for.

The consequences appear years later. No asset record means no criticality rating, so no PM schedule, so maintenance is whatever the controls contractor chooses to do on a site visit. No lifecycle record means no replacement provision, so the day the controller generation goes end of support the money becomes an emergency capital request competing with something visible.

The first practical step is unglamorous: get the system onto the register as a parent with children. The supervisory layer, meaning head end server, operating system, software version and licence entitlement, database engine and workstations. The controller layer, every controller by model, firmware version and plant served, which is the layer that goes obsolete and the one nobody has an inventory of. Field devices as a population by type. Network switches, gateways and protocol converters. And the intangibles: licences, the software maintenance agreement, controller programs, graphics source, commissioning documentation. Those last are assets too, and the ones most often lost. For what the parts actually are, start with the complete guide to building management systems and the points lists and field devices pillar.

The test that exposes the gap

Ask your facilities team a single question: what model and firmware version are the AHU controllers, and when does that generation go end of support? If nobody can answer within a day, the BMS is not being maintained as an asset, it is being attended to on call. That answer is a five minute check on a well run site and a two week investigation on a poorly run one.

2. What BMS maintenance actually consists of

Ask most controls contractors what a maintenance visit covers and you will get an answer about checking the front end and responding to alarms. That is monitoring, not maintenance. Real BMS maintenance has eight workstreams, and a visit that does not touch most of them is not doing the job. This is maintenance of the control system, not the plant it controls: the mechanical side belongs in your HVAC preventive maintenance programme, planned into the same window where the work overlaps.

  • Sensor calibration and verification. Every sensor drifts: temperature slowest, humidity fastest and least forgiving, CO2 depending on the technology. A humidity sensor reading six percent high is not a broken sensor, it is a system conditioning accurately to the wrong setpoint while nobody can see why energy crept up. Verify against a calibrated reference, record before and after values against a stated tolerance, and log the result so you learn which types drift fastest on your site.
  • Actuator and valve stroking. An actuator that has sat in one position for eighteen months may not move freely, and the BMS cannot tell you because it only knows what it commanded. Drive each one through its full range and confirm physically that it follows. You are hunting for the actuator reporting one hundred percent while the valve is stuck at forty, a classic cause of a coil that "does not perform" and gets blamed on the chiller. Include feedback verification and spring return on life safety dampers.
  • Controller battery, firmware and diagnostics. Many controller generations hold their program or clock in battery-backed memory, and those batteries last a few years, so replace on age rather than on failure. Firmware updates fix real defects but occasionally change behaviour: apply them in a planned window with a verified backup taken first. Then read the diagnostics, because a controller restarting every few days is telling you something long before it fails.
  • Server and database housekeeping. Disk space, because trend and alarm history grow without limit unless somebody set retention. Database size and index health, because a supervisory database left to grow for six years is the usual cause of a front end that has become unbearably slow. Service health, certificate expiry, log review, and time synchronisation, which matters more than people expect because schedules, trend timestamps and alarm sequences all depend on it. Patching and network hardening belong in a security programme: see the BMS cybersecurity pillar.
  • Graphics and schedule review. Graphics rot. Plant is replaced and the page still shows the old unit; a zone is reconfigured and the graphic references a point that no longer exists. Walk the graphics tree against current plant, then do the same for every time schedule and calendar exception against how the building is occupied now rather than at handover. The commissioning, graphics and dashboards pillar covers the design side.
  • Alarm review and rationalisation. The highest value hour in the programme, and almost never billed for. Rank the alarm history by frequency and deal with the top twenty. Most will be nuisance: a limit set too tight, an alarm with no delay tripping on every start transient, an alarm on a point that no longer exists. A system generating four hundred alarms a day is generating zero, because nobody reads it.
  • Trend log health. Trends are the evidence base for every energy and comfort investigation you will run, and they fail silently. Check that each one you rely on is logging, at the interval and retention you think, and that none has been stopped by a controller reconfiguration. The failure mode to hunt is the trend dead for eleven months, found at the moment you need last year's data to settle an argument.
  • Backup and restore testing. Its own section below, because it is the task where the gap between what people believe and what is true is widest.

3. A BMS PM schedule you can lift

Below is the task set I would put on a BMS asset record, with frequencies to adjust for criticality and system size. On a small system some collapse into one annual visit. On a hospital, data centre or airport they do not, and quarterly is the floor for the core tasks. The frequency column is the argument you will have with your contractor; the "why" column is how you win it. For building and defending a schedule like this, see the preventive maintenance plans and programs framework.

TaskFrequencyWhy it is on the list
Front end health check: services, disk, database size, backups running, time syncMonthlyCatches the slow-growing failures (disk, database bloat) while they are still cheap
Alarm history review and rationalisation of top nuisance alarmsMonthly or quarterlyRestores trust in the alarm list; the highest value task in the programme
Override register review: list every manual override, justify or release each oneMonthlyOverrides are the main way automation silently stops working
Trend log verification: logging, interval, retention, no dead trendsQuarterlyTrends fail silently and you only notice when you need the history
Controller diagnostics: comms errors, restart counts, memory and CPU loadQuarterlyRestart and comms counters predict controller and network failures
Graphics walk-through against current plant; schedule and calendar reviewSemi-annualBoth diverge from reality with every change, and nobody logs it
Sensor calibration against a reference instrument (sample or full)Annual; humidity and CO2 more oftenDrifted sensors mean controlling accurately to the wrong number
Actuator and valve stroking, full range, with physical confirmationAnnualCommanded position is not actual position; stuck actuators look like plant faults
Controller battery replacement by age, not by failureTypically 3 to 5 yearsA dead battery during an outage can cost you the controller program
Firmware review, applied only in a planned windowAnnual reviewFirmware fixes real defects but can change behaviour
Full backup: programs, graphics source, database, licences, held off siteQuarterly, plus after every changeAn unrecoverable controller program is a rewrite, not a restore
Documented restore test to spare or test hardwareAnnualA backup nobody has restored is a belief, not a backup
Sequence of operation review against as-built intentAnnualRecovers the most energy of any task here; almost nobody does it
Obsolescence and support status check, controllers and softwareAnnualTurns a surprise capital request into a planned lifecycle budget
Spares holding review: controllers, modules, sensors, actuatorsAnnualObsolete controllers get scarce and expensive before they become unavailable

4. The annual sequence of operation review that almost nobody does

If I could add one task to every BMS maintenance contract, it would be this one. Once a year, take the as-built sequences of operation, sit down with the trend data, and check whether the system is still doing what the sequences say. Not whether it is comfortable. Not whether there are alarms. Whether the strategy that was designed, specified and commissioned is still the strategy actually running.

On a building more than three or four years old the answer is almost always no, and the divergence accumulates rather than arriving. A reset schedule disabled during a commissioning dispute and never re-enabled. A free-cooling changeover locked out after a complaint one shoulder season. A chilled water setpoint dropped two degrees to chase a single hot room. A variable speed pump forced to fixed speed to solve a noise issue. Demand-controlled ventilation turned off because a CO2 sensor failed. Each was reasonable on the day. Collectively they undo most of what the control strategy was designed to deliver, and none show up as a fault.

What the review looks like, section by section against the sequence document:

  • Setpoints and resets: are supply air, chilled water, heating water and static pressure resets active and moving as designed, or has everything settled to a fixed value? Trend a week and see whether the reset lines move at all.
  • Occupancy logic: does the plant actually go to its unoccupied state, and does optimum start still work, or was it disabled after one cold morning complaint?
  • Sequencing and staging: are chillers, boilers and pumps staging on the designed logic with the designed lead/lag rotation, or has one machine accumulated four times the run hours of its twin?
  • Simultaneous heating and cooling: trend a reheat coil against its cooling coil. Sustained overlap on the same air stream means the sequence has been defeated somewhere.
  • Interlocks and safeties: freeze, fire, high static and low limit interlocks intact and unbypassed. The one category where a finding is a safety issue, not an energy issue.
  • Deadbands: has somebody narrowed a deadband to chase a complaint and left the plant hunting? Short cycling on a trend is usually a tuning decision, not a fault.

The BMS in HVAC points and sequences pillar sets out what those sequences normally contain. If the as-built documentation cannot be found, and on many sites it cannot, reconstructing it is the first year's deliverable and money well spent: its absence is what makes every future change a guess. Which is why sequence documentation, program ownership and backup obligations belong in the contract from day one, as the BMS specification clauses pillar sets out.

Where the energy actually is

Of everything on the PM table, the sequence and schedule reviews are where energy recovery concentrates, and both are analysis rather than tool tasks: an engineer who understands the control intent, plus half a day with the trend data. Harder to buy than a site visit, which is why it is dropped from the contract first.

5. Service contract structures and what they exclude

BMS service agreements come in roughly three shapes. The names vary by vendor and market but the structures are consistent, and what matters is not the label on the cover but which of the eight workstreams above are actually inside scope, and what happens when something needs replacing.

Scope itemReactive onlyPlanned plus reactiveComprehensive
Call-out and fault attendanceYes, per visitYes, agreed responseYes, agreed response
Labour on repairsCharged hourlyOften includedIncluded
Scheduled preventive visitsNoYes, defined frequencyYes, defined frequency
Sensor calibration and actuator strokingNoSometimes, sample basedUsually, scope stated
Front end and database housekeepingNoSometimesUsually
Backups taken and heldNoSometimesUsually; restore testing rarely
Alarm rationalisationNoRarelySometimes; worth insisting on
Sequence of operation reviewNoNoRarely; usually a separate engagement
Replacement parts (sensors, actuators, modules)ChargedChargedIncluded to stated limits
Controller replacement on failureChargedChargedOften capped or excluded above a value
Software version upgradesNoNoOnly with a software maintenance agreement
Graphics and programming changesChargedCharged, small allowanceHours allowance, then charged
Remote monitoring and diagnosticsNoSometimesUsually, if connectivity exists
Obsolescence and lifecycle adviceNoNoOccasionally; ask for it in writing

The exclusions are where the surprises live, and they are consistent across vendors. Expect, in the small print: power quality events, lightning and surge; anything caused by third party works; network and IT infrastructure outside the controls panels; field wiring and containment; consequential loss and energy cost; anything the manufacturer has declared end of support; software upgrades without a separate maintenance agreement; and work arising from changes the client made themselves. That last matters more than it sounds, because on many sites the client's own team does make changes, and the contract can be read as voiding cover on whatever they touched.

On response times the number matters less than the definition behind it. Ask what starts the clock, whether the response is remote or an attendance on site, and what the resolution target is as distinct from the response target. A four hour remote response with no site attendance commitment is a very different product from four hours on site, and both get written as "four hour response". Remote first is the right model for a BMS, because a large share of faults are diagnosable and often fixable remotely, and an engineer looking at trends within the hour beats a van arriving tomorrow. But pair it with a stated site attendance commitment for faults that need hands, and with a remote access arrangement your security team has actually approved rather than one that quietly exists.

What a comprehensive contract does not buy you

Comprehensive cover protects you against the cost of failure, not against degradation, because degradation is not a failure and nothing in the contract triggers on it. A site can hold the most expensive agreement on the market and still have drifted sensors, defeated sequences, dead trends and forty standing overrides, because none of that generates a fault call. To get degradation addressed you specify the analysis tasks explicitly and pay for the engineering hours. No standard agreement includes them.

6. Who holds the licences, the graphics and the programs

Read this section twice, because the answer determines whether you have a maintainable asset or a relationship you cannot exit. The question is not who maintains the system. It is who possesses the things required to maintain it.

  • Engineering tool licences. Most product lines separate the operator front end from the engineering tool that configures controllers and builds logic. That tool is frequently licensed to the installing contractor, sometimes tied to their dongle or account, sometimes restricted to authorised partners. If it is not licensed to you or to an open pool of competent parties, nobody but the incumbent can change your control logic, and every future change costs whatever they say it does.
  • Source graphics. The compiled graphics running on the front end are not the editable source files behind them. Handovers routinely include the former and not the latter, and without the source a graphics change means a rebuild.
  • Controller programs. Same distinction, higher stakes. On some platforms the compiled program cannot be read back into editable form at all. If the source exists only on one engineer's laptop, you are one hard disk away from a rewrite.
  • Administrative accounts. The site should hold the highest privilege credentials on its own system. It is surprisingly common for that account to be a contractor account the client does not have.
  • The as-built points list with addresses, engineering units, ranges and controller assignments. Without it, integrating or troubleshooting anything starts with a discovery exercise.
  • Documentation: as-built drawings, panel schedules, sequences of operation, commissioning records, network topology and the protocol configuration for every integration.

The consequence is worth being blunt about. Where licences, source and programs sit with a single contractor, you have no competitive market for your own maintenance. You can tender, but any new entrant must either buy into the toolchain or reverse engineer your system, and both get priced into the bid. That is why incumbents on some sites hold the work for fifteen years at prices nobody can benchmark. It is not usually malice, it is the predictable result of a handover nobody scrutinised.

At procurement stage this is fixable with a few clauses and an acceptance checklist. If you inherited the situation, the recovery path is to establish what is missing, request it formally, and price the gap. Open protocols help at the integration layer, and the ASHRAE standards work behind BACnet, documented by BACnet International , is why multi-vendor sites are viable at all. Note the limit though: open protocols make devices interoperable, they do not make a vendor's engineering tool or program source available to you. That is a commercial question, not a technical one. On task content, SFG20 is the usual reference in UK and Gulf specifications.

The question that decides supplier freedom

If we terminated the incumbent tomorrow, could a competent third party take over this system without rebuilding anything? Whatever qualifies that answer is exactly what you need to go and obtain, and the time to ask is while the relationship is good.

7. Obsolescence, end of support and the horizon nobody tracks

A BMS controller generation typically enjoys a long life. It is not unusual for a controls platform to be actively sold for a decade, supported for a decade beyond that, and then genuinely finished. That length is what makes it dangerous: the horizon sits outside normal planning cycles and beyond the tenure of most people involved. The engineer who specified the system has moved on, and whoever inherited it assumes it will keep running because it always has. Obsolescence arrives in stages, and it is worth knowing which one you are in:

  • Current: actively sold, developed and supported.
  • Mature: still sold and supported, development finished, new features going to the successor line. Where most installed bases sit, and a comfortable place to be.
  • End of sale: no longer available new. Support and spares continue, often for years, but every expansion now means used or refurbished hardware. This is where lifecycle planning must start, and the stage most sites sail straight through without noticing.
  • End of support: no firmware fixes, no security fixes, no technical support, spares only from the secondary market. The system still works. It is now a risk you carry rather than an asset you operate.
  • Practically unmaintainable: the engineering tool no longer installs on any supported operating system, the people who knew the platform have left the market, spares have gone. This is where forced replacement happens on the worst possible terms.

A second track moves much faster and catches people out: the server and software side. The supervisory software runs on an operating system and a database engine, both with support horizons measured in years rather than decades, so a controller generation good until the mid 2030s can be stranded years earlier because the head end will not run on anything still supported. Track the software stack separately, because the two will not go end of life together. The BMS software and platforms evaluation pillar is the place to start, and the major vendor comparison covers how the large product lines differ on lifecycle.

The minimum discipline is an annual obsolescence statement: for every controller model, software version and operating system, the current lifecycle stage and the published end of support date where one exists. Request it in writing, keep it with the asset records, take it to budget planning. Two or three years of warning turns a crisis into a project. Worth checking alongside it: what is still in warranty and what that warranty obliges anyone to do, which the warranty management pillar deals with properly.

8. The four upgrade paths, with honest trade-offs

When the obsolescence conversation arrives there are four realistic routes. Vendors have a preferred one, and it is usually not the cheapest.

Head end only. Replace the supervisory software, server and graphics; keep existing controllers and field devices. Cheapest by a wide margin, no plant downtime, and it fixes the fastest-moving obsolescence risk. The honest limits: it does nothing for controller obsolescence, so you are buying years not decades; the new front end may not expose everything the old controllers can do; and you still depend on the old engineering toolchain for any logic change. Good when controllers are mature but healthy, poor when they are already end of support.

Controller replacement, field devices retained. Swap controllers for the current generation, reusing sensors, actuators and wiring where they terminate compatibly. Usually the best ratio of benefit to cost, because it resets the clock on the layer that matters. The trade-offs: logic must be re-engineered rather than copied, so sequences have to be documented or reconstructed first; every reused device carries its remaining life into a new system; and plant comes down zone by zone. Budget more commissioning than the quote assumes, because reconstructing undocumented logic is where these projects overrun.

Staged migration. Install the new system alongside the old, integrate so they present as one front end, migrate area by area over several budget years. The right answer for large estates, hospitals and airports where downtime is genuinely constrained. The costs are often understated: you run two systems throughout, integration between generations sometimes needs gateways, operators live with a split reality, and staged migrations tend to stall halfway when priorities change, leaving a permanent hybrid worse than either endpoint. Commit to a completion date in the business case and defend it.

Rip and replace. New controllers, field devices, wiring where needed, head end, fresh commissioning. Most expensive and most disruptive, and the only option giving a clean, fully documented, fully supported system with a known lifecycle ahead of it. Right more often than budget holders want to hear: where field devices are at end of life anyway, where documentation is irrecoverable, or where building use has changed so much the old strategy is wrong regardless of hardware. If you are rebuilding the logic anyway, much of the saving in the cheaper options evaporates.

The decision driver I would use is not cost, it is documentation. Where sequences, points lists and program source are intact, the cheaper staged options work well because you can carry the intent forward. Where documentation is gone, every option becomes a re-engineering exercise, and the cost gap between controller replacement and full replacement narrows enough that the cleaner outcome often wins.

9. Backup discipline, because an unrecoverable program is a rewrite

Most sites believe they have BMS backups. Fewer have all the components, and very few have tested a restore. A BMS backup is not one file, it is a set, and missing one member can make the rest useless.

A complete set contains: controller programs in editable source form, not only compiled; graphics in source form; the supervisory database including trends, alarms, schedules and user accounts; the configuration and point database; licence files and activation details; the network and integration configuration; and a written note of software and firmware versions, because restoring a program into a controller at a different firmware level does not always behave as expected.

The discipline is straightforward and rarely followed: a scheduled backup at a defined frequency, a further backup before and after every change, at least one copy off site and off the BMS network, several generations of retention so you can go back past a bad change, and an annual documented restore test to spare hardware. That last is the whole point. Backups are a recovery capability, not a storage problem, and an untested backup is a belief. Record who can perform a restore, because on more than a few sites the answer is one individual at one contractor.

What backups cannot recover

A backup restores the program. It does not restore the understanding of why the program is the way it is. Where sequence documentation is missing, a successful restore still leaves you with logic nobody can safely modify, which means every future change is made by trial on a live building. Backups protect against loss. Only documentation protects against paralysis, and they are separate obligations that need separate enforcement.

10. The honest section: BMS degrade through neglect, not failure

Here is the part that does not fit the usual maintenance narrative. The dominant failure mode of a BMS is not hardware failure; controllers are robust and often outlive the plant they control. It is gradual abandonment, and the clinical sign is the accumulation of manual overrides.

The progression is predictable. A complaint arrives and somebody overrides a value to resolve it, intending to investigate properly later. Later does not come. Another complaint, another override. An engineer leaves and the reasoning behind twelve overrides leaves with them. Their replacement finds points forced in ways nobody can explain and reasonably decides the safest thing is to leave them alone and add their own rather than release something they do not understand. Meanwhile the alarm list fills with nuisance, so nobody reads it, so a real alarm goes unseen, so trust falls further. Within a few years the building is run manually through a control system and the automation is a display.

The diagnostic is easy, and I would run it on any site I was asked to assess: pull the full list of manual overrides and forced points, with the date applied and the reason recorded. Three findings are typical. The list is longer than anyone expected, most entries have no recorded reason, and several are years old. Any override that cannot be justified in one sentence by someone currently employed on site is a defect, not a setting.

The counter-discipline is procedural rather than technical:

  • Every override gets a reason and an expiry date, in the system where the platform supports it and in a register where it does not. No anonymous overrides.
  • The override register is reviewed monthly, and every entry is justified again, converted into a proper engineering change, or released. Overrides are temporary by definition; one that should be permanent is a logic change that has not been done.
  • Alarm rationalisation runs continuously, because a credible alarm list is what keeps operators engaged at all.
  • Changes are documented as changes: what was altered, by whom, when and why. Cheapest item on the list, and the one that prevents the knowledge loss causing everything else.
  • The annual sequence review is the backstop, catching drift that slipped past the monthly disciplines.

None of that needs new software or sensors. It needs someone to own the system and a small amount of enforced routine. The uncomfortable implication is that most BMS underperformance is an organisational problem in a technical costume, which is why re-commissioning an existing system so often beats buying a new one, and why replacing a neglected BMS without changing the disciplines around it puts you in the same position in six years with newer hardware.

The idea to walk away with

A BMS is an asset with a service life, and it degrades in a way that almost never announces itself. Sensors drift, actuators stick while reporting perfect position, trends die quietly, graphics diverge from the plant, sequences get defeated one reasonable decision at a time, and overrides accumulate until nobody trusts the automation. None of that is a fault, so none of it triggers a reactive contract, and a comprehensive agreement will not catch it either, because comprehensive covers failure, not decay.

So the programme that works has two halves. A conventional PM schedule for the physical layer: calibration, stroking, batteries, housekeeping, backups with tested restores. And an analytical layer no standard contract includes, which you specify and pay for deliberately: alarm rationalisation, override review, schedule review and the annual sequence of operation review. The second half is where the energy and the reliability come from. Around both, three governance answers: who holds the licences, who holds the source, and when the installed generation stops being supported.

Final thoughts

If you do one thing after reading this, pull the override list and the alarm frequency report. They take an hour between them and will tell you more about the true state of your BMS than any site visit report you have been sent this year. A short list of well-understood overrides and a quiet alarm page means the system is being run. Forty unexplained overrides and four hundred alarms a day means an expensive display panel and a building operated by hand.

And if you are inside the support window with documentation intact, that is the moment to act rather than relax. Obtain the licences and the source while relationships are good, write the analysis tasks into the next renewal rather than hoping they are implied, get the obsolescence statement in writing, and test a restore before you need one. All four are cheap now and expensive later, and later always arrives during an outage, a dispute, or a budget year with no room for surprises.

Disclosure

Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.

Reviewing a BMS service contract or facing controls obsolescence?

Independent advisory on BMS maintenance scope, service contract review, licence and program ownership, obsolescence planning and upgrade path appraisal. 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations. No controls vendor margins, no reseller arrangements.

Book a conversation

Related reading: Building management systems: a complete guide, BMS in HVAC: points and sequences, Points lists and field devices, BMS cybersecurity for connected buildings, BMS specification clauses, Preventive maintenance for HVAC systems.

Muhammad Abbas

CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.

Work with me
MAbbaz.com
© MAbbaz.com