mail@mabbaz.com Abu Dhabi, UAE

Data Centre BMS · Critical Facility Controls · DCIM

BMS for Data Centers and Critical Facilities

In a normal building the controls system exists to keep people comfortable. In a data centre it exists to keep IT load running, and that inverts almost every assumption carried over from commercial work. A practitioner's guide to cooling control, pressure and humidity strategy, electrical monitoring, alarm philosophy, redundancy, and the uncomfortable question of whether the controls system has become the site's biggest single point of failure.

Muhammad Abbas September 25, 2026 ~17 min read

In an office tower the building management system is a convenience layer: schedules, setpoints, an alarm list nobody has cleaned up in three years. In a data centre the same software holds a live thermal and electrical envelope around equipment that throttles within minutes and shuts down within tens of minutes if that envelope slips. Same protocols, same controllers, completely different risk posture. Most serious mistakes I see in critical facility controls come from engineers who are competent at commercial BMS work and have not yet re-examined which habits are safe to bring across.

The message up front: in a critical facility the controls system has two jobs that pull against each other. It must optimise, and it must never be the reason the site goes down. When those conflict, resilience wins. That single rule decides sensor placement, staging logic, alarm philosophy, head-end architecture and change control. If a proposed optimisation cannot survive the loss of the supervisory layer, it does not belong here.

1. Why a critical facility BMS is a different animal

If you are coming from general building controls, the complete guide to building management systems is the baseline this article then changes the constraints on. What differs:

  • The load never goes away. No unoccupied mode, no weekend setback. IT load is continuous and largely insensitive to outside conditions, so almost every energy strategy built around occupancy is irrelevant.
  • Thermal inertia is measured in seconds. An office coasts for an hour on a cooling failure. A well-packed hall with contained aisles and little air volume can leave the allowable envelope in under a minute on total fan loss. Response times that are adequate commercially are not adequate here.
  • A control error does not cause discomfort. It causes throttled compute, then thermal shutdown, and the cost is carried by the business the facility serves rather than the facilities budget.
  • Redundancy is structural. Cooling and power are built N+1 or 2N by design, and the controls layer has to preserve that topology rather than quietly defeating it.
  • Change is governed. A setpoint an engineer would adjust on a Tuesday afternoon elsewhere is here a change record with a written method, a review, a window and a rollback.

The Uptime Institute tier framework is the vocabulary most operators use for how much concurrent maintainability and fault tolerance a site was designed for: see Uptime Institute . The practical point for a controls engineer is blunt. If the mechanical design is concurrently maintainable but your sequence needs both chillers online to hold setpoint, you have silently downgraded the facility.

2. Cooling topologies and what each demands of the controls

Data centre cooling is not one architecture. The equipment type dictates where the control loop closes, which sensors matter, and how fast the system must react. Getting this mapping right is most of the design work.

Topology Where the loop closes Typical density Controls demands
Perimeter CRAC (local DX refrigeration) Unit DX capacity plus fan speed, coordinated across units Up to roughly 5 to 8 kW per rack Team control so units stop fighting; one unit heating while another cools is the classic failure. Compressor staging limits and anti-short-cycle timers.
Perimeter CRAH (chilled water coil) Coil valve plus fan speed against an upstream plant 5 to 15 kW per rack with containment Valve and fan coordination, plant differential pressure reset, supply air temperature control. Plant and unit control must be designed together, not separately.
In-row cooling Close-coupled to the rack row, short air path 15 to 40 kW per rack Fast loops, per-row temperature control, tight fan modulation. Almost no air volume buffer, so response speed and sensor placement dominate.
Rear-door heat exchangers At the rack outlet, passive or fan-assisted water coil 20 to 40 kW per rack Per-door water control, leak detection, and chilled water held above room dew point. Condensation avoidance is a controls constraint, not a plumbing one.
Liquid to the chip (cold plates, CDU) Coolant distribution unit and secondary loop 50 kW per rack and above Secondary loop temperature, flow and pressure control, leak detection, and a clear boundary of responsibility between facility controls and IT-owned equipment.
Free cooling (added to any of the above) Mode arbitration between mechanical and ambient Any Changeover logic with hysteresis, and a hard requirement that mechanical cooling is proven available before ambient mode is trusted. Changeover is the highest-risk sequence in the plant.

Mixing topologies in one hall is normal and legitimate, a CRAH baseline with in-row units on the dense rows, but it multiplies the coordination problem. Two independently tuned systems serving overlapping air volumes will hunt against each other unless one is explicitly subordinate. Decide which system owns the space temperature and give the other a deliberately looser loop.

3. Supply air temperature is the controlled variable, not return air

This is the most consequential control decision in the room and it is still routinely got wrong, usually because the unit controller shipped with a return air sensor and nobody changed the strategy.

Return air control made sense in uncontained rooms with low, evenly distributed load, where return temperature was a fair proxy for room condition. In a contained hall it is misleading. Return temperature is a function of both the cooling supplied and the IT delta T, and IT delta T varies with server fan behaviour, workload and rack population. A row of lightly loaded servers running high fan speeds returns cool air, the unit reads a low return temperature, reduces capacity, and supply temperature at the intake face rises. You have throttled cooling because the load looked light, at the moment a hot spot was forming.

What the equipment cares about is air temperature at the server intake, so control supply air and trim it on worst-case intake:

  • Each unit holds a supply air temperature setpoint measured at its own discharge, keeping the loop local, fast and stable.
  • A supervisory trim adjusts that setpoint from the highest intake temperature across the cold aisle sensor set, so the plant serves the worst rack rather than the average one.
  • The trim has a bounded range and a failsafe value. On loss of supervision the unit reverts to its last known good setpoint or a conservative default and keeps running. It does not stop, and it does not swing fully open or closed.
  • Return temperature is still trended, because a rising return against a stable supply is a useful indicator of increasing load or worsening bypass. Indicator, not control input.
The test I apply on every survey

Ask what the cooling units control to, then ask where that sensor physically is. If the answer is return air, or "the sensor in the unit", the room is controlled on a variable the IT equipment never experiences. Every containment improvement then looks like a cooling failure to the loop, because good containment raises return temperature by design. Teams have torn out containment because it "broke the cooling", when what it broke was a return-air control strategy.

For the underlying grammar of points, sensors and sequences, see the BMS in HVAC controls and sequences pillar and the points lists and field devices pillar.

4. The thermal envelope as the design reference

A controls engineer needs a defensible answer to "what should intake air temperature be?", and it should not be a number remembered from a previous job. The ASHRAE thermal guidelines for data processing environments define equipment classes with recommended and allowable ranges for intake temperature and humidity, and they are the reference the industry actually uses: ASHRAE .

The structure matters more than any specific figure, because figures depend on edition and equipment class. The recommended range is where you should normally sit, benign for reliability across design life. The allowable range is wider, defined per class, permitted by the manufacturer but intended as excursion territory with a time-at-temperature consideration rather than a steady state. And class matters: a hall housing one well-understood class can run warmer than a mixed-tenancy hall where the most sensitive kit sets the limit.

Two consequences for the design. Your supply setpoint should be justified in writing against the recommended range for the classes actually present, because it will be questioned at the next incident review. And your alarm limits should be derived from the envelope rather than picked by feel: a warning as intake leaves the recommended range, a critical as it approaches the allowable limit, with the gap sized so an operator has real time to act.

Raising supply temperature is the most-discussed efficiency lever in the sector and it genuinely works, because warmer chilled water and supply air let the plant run more efficiently and extend economiser hours. It also eats thermal ride-through. A hall at the warm end of the recommended range has less margin before an interruption pushes intake past the allowable limit. That is a fair trade made deliberately with ride-through calculated, and a bad one to drift into by incremental setpoint creep.

5. Chilled water plant: staging that does not spend your redundancy

Where cooling is water-based, the plant is the part that can take the whole site down, and it is where I would concentrate review effort. Topology is normally N+1 or 2N on chillers, pumps and heat rejection, and the question is whether that designed redundancy survives contact with the sequences.

The failure mode that concerns me most is losing redundancy during a stage change. A plant running two of three chillers is N+1, but during the transition to add or drop a machine there is a window where it is not, and a fault inside that window exposes the site. Good staging logic shrinks and manages that window:

  • Prove before you release. Never stop a running machine until the incoming one has proven flow, proven compressor operation and demonstrated it is taking load. Sequencing on a timer alone is how plants trip.
  • Overlap deliberately. Accept over-capacity during a stage change. A few minutes of extra running costs nothing against the exposure of a gap.
  • Hysteresis and minimum run times. Stage-up and stage-down thresholds separated widely enough that normal load variation cannot cause cycling, with minimum run and off timers enforced in the controller rather than assumed from the setpoint spread.
  • Rotate lead duty on a schedule, not on run hours alone. Equalising hours is reasonable, but a rotation firing automatically at an arbitrary moment is an unplanned stage change. Rotate at a known time under supervision.
  • Do not let the standby be a fiction. The standby chiller, pump and tower cell must be exercised and proven to start and take load on a defined cycle. A machine that has not run in six months is a hope, not a redundancy. The mechanics sit in the boiler and chiller preventive maintenance pillar.
  • Differential pressure reset with a floor. Resetting pump head to match demand saves real energy, but the reset needs a floor that guarantees flow to the most hydraulically remote coil. Schemes that chase the least-open valve can starve a remote branch when one valve reports incorrectly.

Where a chilled water buffer or thermal store exists, its job is carrying load across the gap between utility loss and chiller restart on generator power, so its control logic is safety logic, not efficiency logic. Charge state should be monitored and alarmed, and a store that is not fully charged is a reportable condition rather than a trend line.

6. Airflow and pressure control in the white space

Temperature control gets the attention; airflow management determines whether it can work at all. If air bypasses the load or recirculates around it, no setpoint tuning fixes the resulting hot spots. The controllable variable is differential pressure, measured in three distinct places:

  • Underfloor plenum to room. In a raised floor design, holding a modest positive plenum pressure ensures every grille delivers, including those furthest from the units. Too low and remote grilles starve; too high and you waste fan energy and can lift tiles. A slow, stable loop suited to supervisory reset of fan speed.
  • Cold aisle to hot aisle across containment. In a contained hall this is the measurement I would prioritise. A slightly positive cold aisle means the units are delivering at least as much air as the servers draw, so no hot air recirculates to the intakes.
  • Room to surrounding areas. Holding the hall slightly positive against corridors and plant rooms keeps dust and uncontrolled humidity out. Modest, but worth a loop.

This needs honest treatment of sensor quality. The pressures are very small, the sensors drift, and a drifted sensor driving fan speed either starves the hall or runs every fan flat out. On any loop that controls fan speed I would specify redundant sensing, use a median or best-two selection rather than a single input, and alarm on disagreement between sensors rather than only on out-of-range values. A disagreeing pair is a maintenance ticket; a single trusted sensor that has quietly drifted is an incident.

Where controls cannot help

Airflow problems caused by physical layout are not solvable in software. Missing blanking panels, unsealed cable cutouts, grilles in hot aisles, gaps in containment and empty rack positions all bypass or recirculate air. The usual response is to push supply temperature down and fan speeds up until the worst rack is satisfied, which means the whole hall runs cold and over-ventilated to compensate for a handful of holes. A day with blanking panels and brush grommets will outperform months of loop tuning at almost no cost. If a controls engineer tells you the sequences cannot fix your hot spot, that is often the correct and honest answer.

7. Humidity control and the case for loosening it

Humidity control in data centres carries more inherited belief than evidence, and it is one of the few places I would actively argue for doing less than the legacy design assumed. The two genuine concerns are real but bounded: very dry air raises electrostatic discharge risk during handling, and very humid air raises condensation risk and, over long exposure, hygroscopic dust and corrosion concerns. Neither justifies holding relative humidity to within a few percent, which is what a great many legacy CRAC installations were configured to do.

The cost of tight control is substantial and mostly invisible on a dashboard:

  • Simultaneous humidification and dehumidification. Perimeter units with independent local humidity control and no coordination will, reliably, have one humidifying while another dehumidifies in the same room. Each is satisfying its own sensor, the net effect on room humidity is nil, and the energy is entirely wasted. One of the most common and most expensive findings in a critical facility survey.
  • Overcooling to dehumidify. Dehumidification needs a coil below dew point, which means running colder than temperature control requires and often reheating, defeating every warmer-supply-air strategy you implemented.
  • Relative humidity is the wrong variable. It is a function of temperature, so the same air at different points in the hall reads as different relative humidity while moisture content is identical. Units then act on a difference that does not exist. Dew point is common across the space and is what the strategy should use.
  • Humidifiers are maintenance liabilities. They scale, fail, leak and consume attention, in service of a variable most equipment is genuinely indifferent to across a wide band.

What I would recommend: control dew point, not relative humidity. Give the space a single wide dew point band derived from the applicable ASHRAE envelope, manage it at one point in the system rather than at every unit, and disable local humidity control on the individual units so they cannot fight. Alarm on excursions outside the wide band. Hold a hard low limit on chilled water temperature so coils never fall below room dew point, which removes dehumidification-by-accident at source and is also what protects rear-door heat exchangers from condensing.

One caveat. A hall in a hot, humid coastal climate, which describes most of the Gulf, has a genuine latent load whenever outside air enters. That fresh-air path deserves its own dedicated dehumidification rather than being left to the room units. Loosening room humidity control is not the same as ignoring the latent load at the building boundary.

8. Electrical monitoring: UPS, PDU, generator and switchgear

A commercial BMS monitors electrical systems lightly: a run status, a fault contact, perhaps incomer energy. In a critical facility this monitoring is at least as important as the mechanical control, and it is largely monitoring rather than control. What I would expect integrated:

  • UPS. Operating mode, load per module and per system, battery voltage, temperature and estimated autonomy, bypass status, alarm register. Battery temperature deserves attention because battery life is strongly temperature-dependent and a warm battery room is quietly shortening the autonomy you rely on. Mode changes, especially transfers to bypass, are high-priority events whether or not load was affected.
  • PDUs and busway. Per-branch current, voltage, power, energy and phase balance. This tells you whether a rack is approaching its circuit limit and whether the A and B feeds to a dual-corded rack are actually balanced. A rack drawing heavily on one feed is one failure from tripping a breaker during a maintenance transfer, which is a condition you want found before the transfer.
  • Generators. Availability, run status, fuel level and quality indication, coolant and oil condition, start attempts, charger status, and the result of every scheduled exercise run. Fuel level should be a trended value with a low-level alarm tied to required autonomy, not a gauge read on a walk-round.
  • Switchgear and transfer. Breaker positions, source selection, transfer switch position and readiness, protection trip indication, earth fault, plus power quality at the main and UPS input where warranted. Every significant breaker position should appear on a single-line graphic reflecting the real topology, because during an incident the first question is always what is fed from where.

This metering is also what makes efficiency measurement credible. Power usage effectiveness, total facility energy over IT energy, is only as good as the metering behind it: you need facility input energy and IT load energy measured at a defined, documented boundary, ideally UPS output or PDU level rather than estimated. A PUE derived from nameplate assumptions and one incomer meter is a number for a report, not a management tool. The wider measurement discipline is in the energy management systems pillar, the analytics layer that turns this data into detected faults in the FDD and analytics pillar, and the testing regime behind the equipment itself in the electrical preventive maintenance pillar.

One boundary point that avoids trouble: the BMS should monitor the electrical plant and almost never control it. Breaker operation, source transfer and UPS mode changes belong to the electrical protection systems and to trained operators following written procedures. A BMS write command into switchgear carries real risk for very little upside. Read everything, write nothing, is a defensible default.

9. Alarm philosophy when a 3am call is expensive and a missed one is worse

Alarm design here is genuinely hard because both failure directions cost. Too many alarms and the on-call engineer learns to dismiss them, which is how a real event gets missed at 4am. Too few, or thresholds too wide, and the first indication of trouble arrives from the IT side. Most sites I look at are firmly in the first category, with alarm lists in the hundreds and a team that has stopped reading them. The principles I would hold to:

  • Every alarm has a defined response. If nobody can say what action it requires, it is a status point or a trend. Move it. This one rule typically removes half of a bloated list.
  • Three tiers, not ten. Critical: IT load is at risk now, act now. High: redundancy lost or a protective condition exists, act within a defined time. Advisory: pick it up in normal hours. Three clear tiers beat a five-level scheme everyone interprets differently.
  • Loss of redundancy is a first-class alarm. The one most commercial-derived designs miss. A chiller failing while its standby picks up load has not affected temperature and will raise no temperature alarm, but the site has moved from N+1 to N and is one fault from an outage. That must alarm loudly on its own merits and be visible as a state on the graphics, not just an event that scrolled past.
  • Suppress the cascade, not the root. One chiller trip should produce one alarm, not forty. Parent and child relationships, first-out logic and shelving during known maintenance states keep the list readable. Shelving must be time-limited, automatically reinstated, and visible as shelved.
  • Delays and deadbands on everything analogue. A point that momentarily crosses a limit during a stage change should page nobody. An on-delay matched to the real response time of the system removes most nuisance alarms without meaningfully delaying detection of a genuine excursion.
  • Route by tier and review monthly. Critical alarms wake someone; advisory alarms go to a daylight queue. Then look at the most frequent alarms each month and either fix the condition or fix the alarm. An alarm that has fired two hundred times and never required action is actively harmful.

The measure of a healthy critical facility alarm system is boring: on a normal day the active list is short enough that an operator reads all of it, and every entry is something somebody intends to act on. If your list fails that test, rationalising it is higher-value work than most optimisation projects.

10. Controls resilience: not being the single point of failure

This is the section I would ask any critical facility controls designer to defend in detail, because a controls system layered over properly redundant plant can quietly reintroduce the single point of failure the whole design was built to eliminate.

Failure mode Consequence if unaddressed Design response
Single supervisory server or head-end fails Loss of visibility, alarming and any logic living at supervisory level; operators blind during an event Dual head-ends, ideally in separate rooms on separate power paths, with independent alarm routing. No logic the plant depends on lives only at supervisory level.
Supervisory link to field controllers lost Units revert to an undefined state, or hold a stale trimmed setpoint that no longer suits the load Failsafe local control: each controller holds a defined safe setpoint and runs its local loop independently. Specified explicitly and tested, not assumed from the datasheet.
One controller governs multiple redundant units A single controller failure removes duty and standby together; redundancy defeated in software Controller independence: redundant equipment gets independent controllers on independent power and network segments. One controller never owns both halves of a redundancy set.
Network is a single ring or one switch stack One fault isolates a whole plant area from supervision and from peer-to-peer coordination Segmented network with redundant paths, controls network separate from corporate IT, and peer-to-peer dependencies kept within a segment where possible.
Controls panels on non-essential power Controls die on utility loss at the exact moment coordinated plant restart is needed Every panel, sensor power supply, network switch and head-end on UPS-backed and generator-backed supply. Verify on the actual distribution schedule, do not take it on trust.
Shared sensor feeding redundant units One failed sensor drives both duty and standby to the same wrong output Independent sensing per redundant path, with cross-checking and disagreement alarming. Voted or averaged inputs where one value must be shared.
Free cooling changeover fails mid-transition Neither mechanical nor ambient cooling fully established while load continues Prove mechanical cooling available before releasing it, overlap the modes, and fail to mechanical on any changeover fault.
Remote access path compromised Unauthorised change to live setpoints or sequences in a critical facility Segmentation, no direct internet exposure, brokered and logged remote access, role-based permissions with write access tightly held.
No current record of the as-built logic Nobody can safely modify or restore the system; a controller failure becomes reverse engineering under pressure Version-controlled configuration backups held off the controls network, tested restores, and a maintained sequence-of-operation document matching what is actually loaded.

That last row is the one most often skipped. A failed controller is a two-hour job if you hold a tested backup and a current sequence document. Without them it is a multi-day reconstruction, by whoever is available, on a live critical site. Backups and documentation are resilience infrastructure, not administration. Network exposure deserves its own treatment too, since controls systems in critical facilities are a recognised target and remote access is usually the weakest link: see the BMS cybersecurity pillar.

11. Change control on anything touching a live sequence

A setpoint change in a data centre is a change to the live protection envelope around running compute, and it should be governed accordingly. This discipline is what most distinguishes a mature critical facility from a competent commercial one. Around any controls modification, from a setpoint tweak to a sequence rewrite, I would expect:

  • A written statement of what changes and why, with the current value, the proposed value and the reasoning. "Improve efficiency" is not a reason; "raise supply setpoint one degree to extend economiser hours, ride-through impact assessed" is.
  • A risk assessment against the thermal and electrical envelope, stating the worst case if the change misbehaves and whether redundancy is reduced at any point during implementation.
  • A rollback plan with the original configuration captured first, so reverting is a restore rather than a reconstruction from memory.
  • A defined window and a defined observer. Changes go in while the site is attended and the plant is normal, never during an activity that has already reduced redundancy, and someone watches the affected trends afterwards.
  • Write access restricted by role. Most people who need the system only need to read it. The accounts that can modify a live sequence should be few, named and logged.
  • A change log that is actually maintained. Most unexplained control behaviour in an older facility is an undocumented change from years ago.

The parallel maintenance-side discipline, method statements, permits, approval routing, vendor coordination and spares, belongs to the work management system rather than the BMS, and is covered in the CMMS for data centres pillar. The two need to interlock: a controls change should raise a work record, and a work record that alters plant availability should be visible to whoever is watching the controls.

12. Where BMS ends and DCIM begins

Data centre infrastructure management and building management overlap enough to cause duplicated data and arguments about ownership. Drawing the boundary at design stage saves a great deal of friction. The clean framing: the BMS owns the facility systems that keep the environment inside its envelope, and DCIM owns the IT estate and the space, power and cooling capacity allocated to it.

  • BMS territory: cooling plant and unit control, chilled water, airflow and pressure, humidity, mechanical sequencing and staging, environmental sensing, plant alarming, and monitoring of electrical infrastructure up to the point of delivery.
  • DCIM territory: rack and asset inventory, rack elevations, port and cable management, per-rack power and space capacity planning, provisioning workflow for IT deployments, and forecasting where the next rack can go.
  • The overlap: per-rack power draw, per-rack intake temperature, and space cooling capacity. Needed by both, and exactly where duplication happens.

For the overlap the rule that works is one system of record per data point, with the other consuming it. Rack-level power from intelligent PDUs is best collected once and shared rather than wired into two systems. Rack environmental sensing can sit on either side, but it should sit on one. Where DCIM needs environmental context for a capacity decision it consumes it from the BMS by interface rather than deploying a parallel sensor network, and where the BMS needs IT load for PUE it takes it from metering both parties agree on.

The integration that pays for itself

The genuinely valuable DCIM to BMS link is not a merged dashboard, it is capacity truth. DCIM knows a deployment is planned for a specific row; the BMS knows the real cooling capacity and current airflow in that row. Connecting the two lets a capacity decision rest on measured conditions rather than a spreadsheet of design allowances. Almost every high-density hot spot I have seen was created by a rack placed where the design said there was capacity and the physical airflow said there was not.

Both systems also need to feed the operator interface without creating three competing versions of the truth. The discipline for that, a sane graphics hierarchy, one clear situational view, drill-down that follows the plant, is in the commissioning, graphics and operator dashboards pillar.

13. Commissioning: integrated systems testing and load banks

Commissioning a critical facility control system differs from a commercial one in that you must test failure, not just function. A system that behaves correctly in normal operation has demonstrated almost nothing about the scenarios it exists for. The stages, in order:

  • Point-to-point verification. Every sensor, actuator, status and command proven from field device to graphic, with sensors verified against a calibrated reference at the installed location rather than accepted because a plausible number appeared.
  • Sequence verification per subsystem. Each unit, plant and loop demonstrated against the written sequence including setpoints, timers, interlocks and alarm limits. Where behaviour differs from the document, reconcile document and logic before sign-off in whichever direction is correct.
  • Load bank testing. Resistive load banks simulate IT load so cooling and electrical systems are proven at design capacity before live equipment is at risk. Test partial loads as well as full: staging and efficiency problems show up at part load, and a plant that is stable at full capacity can hunt badly at 30 percent, which is where a new facility actually operates for its first years.
  • Integrated systems testing. The stage that matters most and gets cut most often. Scripted failure scenarios run end to end with load banks in place: utility failure and generator start with the cooling plant restarting in sequence; loss of a chiller and timed standby pick-up; loss of a unit within a row; loss of the head-end, confirming units continue on local control at failsafe setpoints; network segment failure; free cooling changeover in both directions and a fault induced mid-changeover; fire alarm interface, shutdown and the restart sequence afterwards.
  • Thermal ride-through measurement. With load banks running, remove cooling and measure how long intake temperatures take to reach the allowable limit. This is the single most operationally useful number the facility owns, and it is measured rather than calculated. Every response-time decision should reference it.
  • Alarm and notification verification. Each alarm triggered deliberately, confirmed to annunciate at the right tier and to reach the intended recipient by the intended route. Untested notification paths fail silently, on the night you needed them.
The honest limitation of all this

Integrated systems testing is expensive, takes weeks of programme, needs load banks and temporary power, and is the easiest thing to compress when a handover date slips. It is also the only point in the facility's life when you can safely break things to see what happens. Once live load is in the hall, every one of those tests becomes either impossible or a governed high-risk activity with a written method and a nervous audience. A facility that skipped integrated testing does not have an untested controls system; it has one that will be tested for the first time by a real failure, at an unplanned moment, with load on. That is the trade made when the testing budget is cut, and it should be made explicitly by someone senior enough to own it.

The idea to walk away with

A data centre BMS is not a commercial BMS with tighter setpoints. It is a different design problem in which the controls system is part of the critical infrastructure rather than a layer above it. That reframing drives every specific recommendation here: control to supply air because that is what the load experiences, derive limits from the thermal envelope rather than habit, stage plant so redundancy is never spent in a transition, loosen humidity control because tight bands cost real money for marginal benefit, monitor electrical thoroughly and control it almost never, keep the alarm list short enough to be read, and architect the controls so losing any part of them degrades the facility gracefully instead of taking it down.

One question worth carrying into every design review: if this controls system failed entirely right now, would the plant keep the hall inside its envelope? If the answer is no, the controls have become a single point of failure in a facility built specifically not to have one, and that is more urgent than any efficiency opportunity on the list.

Final thoughts

The pattern I see most often is not incompetence, it is inherited commercial habits nobody re-examined when the stakes changed. Return air control, per-unit humidity setpoints, a single head-end, an alarm list grown by accretion, a staging sequence tuned for efficiency by someone who did not know the site was meant to be concurrently maintainable. Each is defensible in an office building and each is a real exposure in a data centre. Finding them is mostly a matter of asking what the controls are measuring, where the sensor is, and what happens when the layer above it fails.

If you are reviewing an existing site, I would look at three things in this order. What do the cooling units control to, and where is that sensor. Does loss of redundancy raise an alarm on its own merits. And does the plant hold the hall inside its envelope with the supervisory layer switched off. Those three answers will tell you more about real resilience than a week spent with the energy trends.

Disclosure

Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.

Reviewing controls resilience on a critical facility?

Independent advisory on critical facility controls architecture, cooling and plant sequencing review, alarm rationalisation, BMS to DCIM boundary definition and integrated systems testing scope. 22+ years across enterprise CMMS, EAM, CAFM and facility systems integration. No controls vendor margins, no reseller arrangements.

Book a conversation

Related reading: Building management systems: a complete guide, BMS in HVAC: controls, points and sequences, Points lists and field devices, BMS commissioning, graphics and operator dashboards, BMS cybersecurity for connected buildings, CMMS for data centres.

Muhammad Abbas

CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.

Work with me
MAbbaz.com
© MAbbaz.com