mail@mabbaz.com Abu Dhabi, UAE

FDD · Building Analytics · Energy Optimisation

BMS Energy Optimisation, FDD and Analytics

Automated fault detection and diagnostics is the most immediately useful analytics layer you can put on a building management system. It finds the simultaneous heating and cooling, the economiser that stopped economising, the override left on in 2019. It also produces several hundred faults on its first run, and the programme lives or dies on whether somebody prioritises that list and owns acting on it.

Muhammad Abbas September 25, 2026 ~22 min read

Most buildings do not have an energy problem, they have a controls problem that shows up on the energy bill. The plant is competent, the air handling units were commissioned properly, the sequences are sound. Then years of overrides, failed sensors, stuck actuators and well-meaning manual interventions accumulate, and the building quietly drifts into spending more energy than it needs to while nobody notices, because nothing alarms. Fault detection and diagnostics is the discipline of noticing. It reads the trend data the BMS already collects, applies engineering logic, and tells you which plant is misbehaving and roughly what that costs. It is one of the highest-return analytics investments in building operations, and also the one most likely to end up as shelfware, for reasons that have little to do with the software.

The message up front: FDD is a findings engine, not a savings engine. The software reliably tells you what is wrong. It fixes nothing, and it will not prioritise unless you configure it to. The buildings where FDD delivers are the ones where a named engineer owns a ranked weekly list of the top ten or twenty faults, closes them through the CMMS, and can show the trend data flattening afterwards. The rest of this guide is detail around that one organisational fact.

1. What fault detection and diagnostics actually is

A BMS alarms on a state: a sensor out of range, a pump that failed to prove, a space temperature beyond a deadband. Those alarms are binary, instantaneous, and tell you something has already broken. FDD sits a layer above. It reads trend histories across multiple points on the same piece of plant, applies logic encoding how that plant should behave, and identifies conditions where nothing has technically failed but the equipment is operating wrongly.

That distinction is where most of the value lives. A chilled water valve stuck part-open does not alarm. Every point is in range and the space temperature is probably fine, because the control loop compensates elsewhere. But the air handling unit is now heating and cooling the same air stream, burning energy in two directions at once, and will do so indefinitely until somebody walks past with a thermal camera or a rule notices both valves have been open simultaneously for most of the year's occupied hours.

So FDD answers three questions in sequence. Detection: is this equipment operating outside the envelope it should be? Diagnosis: which fault explains the pattern, and what is the likely physical cause? Evaluation: what is it costing in energy, comfort or equipment life, and therefore where does it sit on the queue? A tool that only does the first is an alarm engine with better maths, and that is what produces the wall of alerts that kills programmes. For how analytics sits on top of a controls estate, see the smart buildings maturity path, which owns the data-layer discussion I am not repeating here.

2. Rule-based versus data-driven, and why rule-based dominates buildings

There are two broad families of FDD approach, and the industry conversation is skewed by machine learning marketing in a way that does not reflect what is actually deployed in buildings.

Rule-based FDD encodes engineering knowledge as explicit logic. If the outside air temperature is below the economiser changeover point, and there is a call for cooling, and the outside air damper is still near minimum, then the economiser is not economising. The rule is written once by someone who understands the sequence, applies to every unit with that sequence, and its output is a sentence a mechanical engineer can act on. Data-driven FDD learns normal behaviour from historical data, statistically or with machine learning, and flags deviation. It needs no explicit knowledge of the sequence, which is its attraction, and it can catch faults nobody wrote a rule for.

In practice, rule-based approaches carry the overwhelming majority of deployed building FDD, and I think that is the correct engineering answer rather than industry conservatism. Three reasons:

  • The physics is known. An air handling unit is not a mystery. We know the psychrometrics, we know the control sequence because we wrote it, we know what a working economiser looks like. When you already hold a first-principles model of the equipment, learning that model statistically from noisy trend data is a strange thing to do.
  • The faults are enumerable. HVAC plant fails in a finite, well-catalogued number of ways: stuck dampers and valves, sensor drift, leaking coils, fouled filters, schedule overrides, short cycling, starved terminals. The list fits on two pages. That is unlike fraud detection, where the adversary invents new patterns and enumeration is impossible.
  • Explainability decides whether anyone acts. The one that matters most and gets discussed least. FDD output is consumed by an engineer who must go to a piece of plant and do something. "AHU-04 heating and cooling valves simultaneously open most of occupied hours, check heating valve actuator linkage" gets acted on. "AHU-04 anomaly score 0.83" does not, and after a few weeks of unexplained scores the engineer stops opening the tool.

Data-driven methods earn their place as a complement: catching novel degradation no rule covers, flagging whole-building consumption that has drifted without any single unit looking faulty, and pattern work on plant whose sequences are undocumented. The pattern I would recommend is a rule library carrying the operational load with learned detection layered on top. For the asset-side equivalent, see AI anomaly detection and early fault warning, which covers anomaly detection in a rotating-equipment context.

The explainability test

Before you buy any FDD tool, ask the vendor to show you a fault as the technician will see it. If the output does not name the equipment, name the suspected physical cause, and quantify the impact, the tool will not survive contact with a busy maintenance team. Every fault should read like a work order waiting to be raised.

3. The classic fault library and what each one costs you

You are not starting from a blank page. The same faults recur across every commercial estate I have looked at, and a competent rule library ships with most of them. Knowing the library matters even if you never write a rule, because it tells you which trend points you need and what to expect on the first run.

Fault How it is detected Typical cause Impact
Simultaneous heating and cooling Heating or reheat output and cooling output both above threshold in the same interval Stuck or leaking valve, failed actuator linkage, competing setpoints, reheat misconfigured Very high energy. You pay twice for the same air. Often the largest fault in an estate
Economiser not economising Cooling call present, outside air favourable, damper still near minimum Failed damper actuator, disabled economiser logic, outside or mixed air sensor wrong High energy. Mechanical cooling running when free cooling was available
Economiser stuck open Damper open while outside air is hotter or more humid than return air Actuator failed open, linkage disconnected, control reversed after a panel change High energy plus humidity problems. Cooling load imported from outside
Stuck damper or valve Commanded position varies widely while the measured flow or temperature difference does not respond Seized actuator, broken linkage, debris, failed positioner, pneumatic air loss Medium to high. The loop cannot control, other plant compensates, complaints follow
Sensor drift or failure Flatlined, physically implausible, or persistently divergent from a neighbouring sensor Ageing element, lost calibration, sensor in the wrong location, wiring or module fault Insidious. Every downstream control decision is wrong, and it corrupts FDD and metering both
Schedule override left on Plant running in scheduled unoccupied hours, or a point left in manual Temporary override for an event or complaint, never released. The commonest fault anywhere High energy, zero capital cost. Pure waste, hours to remedy
Short cycling Starts per hour above an equipment-specific limit, or runtimes below a minimum Oversized plant, narrow deadband, staging thresholds wrong, load far below design Equipment life and energy. Starts wear compressors and motors, not run hours
Excessive setpoint deviation Controlled variable persistently outside a tolerance band during occupancy Undersized or starved coil, wrong valve authority, poor loop tuning, upstream supply not met Comfort first, energy second. The fault your occupants already report as complaints
Terminal unit starvation VAV or FCU at maximum flow request yet not meeting setpoint, often clustered on one branch Static pressure setpoint too low, balancing damper closed, obstruction, undersized branch Comfort. Often misdiagnosed as plant capacity and fixed by raising fan pressure everywhere
Reset sequence not implemented Supply air, chilled water or duct static held fixed with no reset against load Reset never commissioned, or disabled during a complaint and left off Medium to high energy, no hardware cost to fix
Valve or coil leak-through Temperature change across a coil while the valve is commanded fully closed Worn seat, debris, wrong valve selection, actuator not fully closing Medium energy, continuous. The quiet version of simultaneous heating and cooling

Two observations. First, an unusually large share of high-value faults cost nothing to fix: releasing an override, re-enabling a reset, correcting a minimum damper position. Engineering hours, not capital. Second, the faults that generate the most occupant complaints, setpoint deviation and terminal starvation, are rarely the ones costing the most energy, which is exactly why you need an explicit prioritisation scheme rather than following the complaint queue. To read any of these rules you need to know what points exist and what the sequences do, the ground covered in the BMS in HVAC controls, points and sequences guide and the points lists and field devices guide.

4. Fault prioritisation: giving the engineer a ranked list

Read this section twice, because it is the difference between a programme that works and one that is quietly abandoned. FDD tools detect faults in the hundreds. A facilities team can act on a handful per week. So the entire practical value of the system is determined by how well it orders that list, and prioritisation is a configuration decision, not a feature you get for free. The model I would recommend scores each fault on four factors and ranks on the combination:

  • Energy impact: estimated annualised energy cost, from the affected load, the magnitude of the deviation and how long the fault has persisted. Order of magnitude is enough.
  • Comfort and compliance impact: how many occupied spaces are affected, whether a ventilation, humidity or temperature requirement is breached, whether a critical space such as a data room, laboratory or clinical area is involved.
  • Asset risk: whether the fault is degrading the equipment itself, short cycling and continuous full-speed operation being the obvious cases.
  • Cost and ease of remedy: an override release is minutes, a stuck damper on a roof unit is a scheduled visit, a rebalancing exercise is a project. High impact plus low effort goes first, always.
Priority Profile Typical faults Action and route
P1 High energy or compliance impact, remedy is configuration only Overrides left on, disabled reset sequences, economiser logic switched off, excess minimum outside air Fix this week from the controls workstation. No work order beyond a logged change record
P2 High energy impact, remedy needs a site visit and a part Simultaneous heating and cooling from a failed valve, stuck economiser damper, coil leak-through Corrective work order in the CMMS with the fault evidence attached, targeted this month
P3 Comfort or compliance impact in occupied or critical space Excessive setpoint deviation, terminal starvation on a branch, critical-space humidity excursion Work order, but investigate as a system. Check static pressure and balancing before touching plant
P4 Asset risk, moderate energy Short cycling, continuous full-speed fan or pump operation, wrong staging thresholds Engineering review of sequences and sizing. Often a design correction, not a repair
P5 Data integrity: blocks other diagnosis Sensor drift, flatlined points, implausible readings, missing trends Fix early despite low energy value, because every fault downstream of a bad sensor is unreliable
P6 Low impact, informational Brief transients, minor deadband excursions, faults on decommissioned or seasonal plant Suppress or aggregate. Do not put these in front of the engineer at all

Note where sensor faults sit. Their direct energy cost is small, so an energy-weighted ranking buries them, but they poison everything above them, which is why I pull data-integrity faults forward out of sequence. And note P6: willingness to suppress faults is a sign of a mature programme, not a lazy one. Twelve real items that get closed beats four hundred that get ignored.

The rule I would hold to

The engineer never sees the full fault list. They see the top ten or twenty for this week, already ranked, each with the equipment, the suspected cause, the estimated impact and the trend chart that proves it. Everything else stays in the tool until it rises. If you cannot deliver that view, you are not ready to turn FDD on across the estate.

5. Continuous commissioning as the operating model

Traditional commissioning is an event. The building is handed over, systems are proved against design intent, a witness sheet is signed, and the folder goes on a shelf. Everyone knows what happens next: performance decays from that day forward. Sensors drift, actuators seize, sequences get modified for a complaint and never restored, tenants change, the building gets fitted out differently from how it was balanced. Within a few years it bears only a partial resemblance to the building that was commissioned.

Continuous commissioning treats that decay as the normal condition and puts a permanent process against it. FDD is what makes it practical, because the alternative, sending engineers to re-verify sequences manually on a cycle, is too expensive to do often. The operating model:

  • Weekly: a named engineer reviews the ranked list, closes configuration-only items directly, raises work orders for the rest, and suppresses anything that is not a real fault. Thirty to sixty minutes when the list is properly prioritised.
  • Monthly: closure rate, recurring faults, and faults closed that have reappeared. A fault that keeps coming back is either a rule tuned wrongly or a root cause never addressed.
  • Quarterly: re-verify a rotating sample of sequences against design intent, review setpoints and schedules against actual occupancy, re-tune noisy thresholds.
  • Annually: performance review against baseline, a sequence audit on plant that changed, and a fault-library refresh.
  • On every change: log any modification to a sequence, setpoint or schedule, with a reason and an expiry date if temporary. This one discipline eliminates the largest fault category in most estates.

That last point deserves emphasis. Override faults dominate the list in most buildings not because anyone is careless but because nothing forces a temporary change to expire. If your controls workstation supports timed overrides, use them exclusively; if not, keep a change register and review it monthly. That is a governance fix, and the cheapest energy saving available to any facilities team. For the operator-facing side, how graphics need to be built so a technician can verify what a fault is telling them, see BMS commissioning, graphics and operator dashboards.

6. Measurement and verification: why claimed savings need a baseline

A conversation I have had more than once. A team runs FDD for a year, closes a few hundred faults, and the tool reports a large cumulative avoided-energy figure. Finance asks whether actual consumption fell by that amount. It did not, nobody can explain the gap, and the programme's credibility takes a hit it does not deserve. The gap has ordinary causes. FDD savings estimates are engineering calculations of what a fault was costing, summed per fault. They do not account for interaction between faults, a different weather year, risen occupancy, new IT load, or the fact that fixing a starved terminal unit can legitimately increase consumption while improving comfort. That does not make them useless. It makes them estimates, to be reconciled against measured consumption rather than reported as though measured. What I would recommend, in order:

  • Establish the baseline before you fix anything. At minimum twelve months of whole-building and major-submeter consumption, with weather and occupancy data for the same period. Fix faults first and you have lost the ability to prove anything.
  • Adjust for the independent variables. Degree days, occupancy hours, connected load. A building that consumed less because the summer was mild did not save energy. The recognised framework for adjusted comparison is the International Performance Measurement and Verification Protocol maintained by the Efficiency Valuation Organization .
  • Keep FDD estimates as estimates, clearly labelled. Use them for prioritisation, not in a savings report as though they were meter readings.
  • Verify fixes at the equipment level and report both numbers. The most defensible evidence is trend data: valves that overlapped for most of occupied hours before the fix and not at all after. Adjusted whole-building performance is the headline, fault-level verification the evidence for what drove it.

The metering, normalisation and management-system side of this, submetering design, weather normalisation and ISO 50001, belongs to the energy management systems for buildings guide, and I would read that alongside this one rather than treat FDD as a standalone energy programme. FDD finds faults; the EMS tells you whether the building actually got better.

Where the savings claim breaks down

Summing per-fault estimates double-counts, because faults interact. A stuck heating valve and a failed economiser on the same air handling unit are both charged the same wasted cooling energy; fix one and the other's estimated saving shrinks. Any report that adds up hundreds of independent estimates and presents the total as annual savings is overstating, sometimes by a large factor. Use the estimates to rank, use the meter to report.

7. The data prerequisites that decide whether FDD is even possible

FDD is arithmetic on trend data. If the data is not there, at the right interval, for long enough, with names you can interpret, no software will help. This assessment belongs before procurement, and it is the step most often skipped, which is why projects stall at integration.

Points. Each rule needs a specific set of points and cannot run without all of them. A simultaneous heating and cooling rule needs both valve outputs. An economiser rule needs outside air temperature, mixed or supply air temperature, damper position and a cooling demand signal. Expect gaps: damper and valve position feedback are the two most commonly absent, and without feedback you are inferring from command, which is exactly the assumption a stuck actuator violates.

Trend intervals. Fifteen minutes is workable for most energy rules; five is better and what I would specify for new work. Hourly averages destroy the short cycling and transient rules, and will mask simultaneous heating and cooling if the events alternate within the hour. Change-of-value trending is storage-efficient but complicates analysis, so check the tool handles irregular intervals.

Retention. Thirteen months minimum, so any month compares against the same month last year. Many BMS configurations retain weeks at the controller and never archive.

Naming and tagging. This is where the real project cost sits. A rule written against "the cooling valve position of an air handling unit" needs to know, for every unit, which point that is. If your point names are the legacy of three contractors across fifteen years, somebody has to map them, and that mapping is usually the largest line item in a deployment. A consistent semantic tagging model, required in new controls specifications, is what makes rules portable across the estate instead of hand-configured per unit. Insist on it from now on, because retrofitting costs several times more than specifying. The governing reference for HVAC control intent, including the guideline series on commissioning and control sequences, sits with ASHRAE .

A pre-procurement checklist you can lift directly:

1. Export the full points list for a representative sample of plant.
2. Take your twenty highest-value rules and list the points each requires.
3. Mark each point present, present-but-unreliable, or absent.
4. Check trend interval and retention for every present point.
5. Count how many of the twenty rules can run today, unmodified.
6. Price the gap: sensors, feedback, trend configuration, archive, naming and mapping.
7. Only then talk to vendors, with that gap cost in the business case.

8. The organisational part: somebody must own acting on findings

Every failed FDD deployment I am aware of failed here, not in the technology, and the pattern is consistent enough to describe as a sequence. The tool is procured with energy savings in the business case. It is integrated, the rules are enabled, and it produces a large number of findings. The findings go to a shared inbox or a dashboard that nobody's job description mentions. The facilities team, already fully committed to reactive work and statutory compliance, has no capacity to absorb a new queue. Some faults get fixed opportunistically, the list grows, the licence is renewed once out of sunk-cost reluctance, and the year after that it is not. What I would put in place instead:

  • A named owner with allocated time. Not a committee, not "the energy team", one named engineer with defined hours per week. Half a day a week for a mid-sized estate is a realistic start. If nobody's time is allocated, nobody's time will be spent.
  • Findings route into the CMMS, not into email. A fault that becomes a work order enters the system the technician already works in, gets planned, assigned and closed, and leaves a history. A fault that stays in the analytics tool competes with the work order queue and loses. This is the most important technical requirement after prioritisation, and it is routinely treated as a phase two item that never arrives.
  • A closure loop back to the tool. When the work order closes, mark the fault resolved and check the trend data confirms it went away. Without that loop you cannot distinguish a fixed fault from an ignored one.
  • Findings in the contract. If maintenance is outsourced, the service agreement must say who investigates findings, within what response time, and how closure is evidenced. Otherwise the contractor's scope legitimately excludes the queue you just created.
  • A metric that is reported. Closure rate, median age of open high-priority faults, recurrence rate. Put them in the monthly FM report beside the reactive and PM metrics, where the FM KPI framework already sets the pattern. Anything unreported drifts.

The sequencing implication, bluntly: if you cannot secure an owner with allocated hours and a route into the CMMS, do not buy FDD yet. The software is not the constraint, and buying it earlier accelerates nothing. It just starts the shelfware clock.

9. A deployment sequence that survives the first year

  • Data readiness on one pilot building. Fix missing feedback, trend configuration, archive and naming for that building only.
  • Baseline. Twelve months of adjusted consumption, or what you have with the confidence stated honestly.
  • A narrow rule set. Ten to fifteen rules, the high-value energy ones plus sensor integrity. A narrow set produces a list the team can clear, which builds the habit.
  • Tune before you widen. Six to eight weeks removing false positives. Every one you leave in teaches the engineer to distrust the tool.
  • Close the CMMS loop. Prove end to end that a fault becomes a work order, gets closed, and that trend data confirms resolution.
  • Widen the rule set, then the estate. In that order, because rule tuning transfers between buildings and process maturity does not transfer if it never existed.

The temptation, usually from whoever signed the estate-wide licence, is to enable everything everywhere on day one. That produces the fault wall on every site at once, no capacity to tune and no habit. Narrow and deep beats broad and thin, for the same reason it does in predictive maintenance.

10. The honest section: the first run produces hundreds of faults

This is the part vendors do not lead with, and the part you most need to plan for. The first FDD run on any building operating for a few years produces a very large number of faults. Hundreds is normal for one mid-sized commercial building. This is not a defect in the tool, and not a sign your building is badly run. It is several years of invisible drift becoming visible at once, plus a meaningful proportion of false positives from rules not yet tuned to your plant. Three things follow, and they are the practical heart of this guide.

First, most of the value is in the first twenty. Fault impact is severely skewed. A small number of faults, typically the overrides, the simultaneous heating and cooling on the largest air handling units, the disabled resets and the failed economisers on the biggest loads, account for most recoverable energy. The long tail of minor deviations on small terminal units is real but individually worth very little. Work the list by impact and you capture most of the available benefit in the first few months; work it in the order the tool output it and you spend that time on trivia.

Second, much of the first list is not actionable. Some are false positives from untuned thresholds. Some are real but relate to plant that is decommissioned, seasonally idle, or due for replacement. Some are one root cause expressed as six rule hits on the same unit. Triage and consolidation before the list reaches anyone is mandatory work, not tidying, and it takes an experienced engineer real days per building.

Third, and this is what kills programmes: hand over an unprioritised list and the programme dies. An engineer presented with four hundred undifferentiated faults makes the rational decision, which is to conclude the tool is noise and stop opening it. That is not a motivation or training problem. It is a legitimate response to an unusable queue from someone who already has a full one. The failure is upstream, in whoever chose to expose raw output instead of a ranked short list.

What FDD cannot do

It cannot find faults in equipment that is not trended, and cannot see a stuck actuator with no position feedback. It cannot distinguish a faulty sequence from a badly designed one, so it will flag the symptom of an undersized coil forever without telling you that is what it is. It cannot fix a building that is oversized or poorly zoned, and it will generate persistent unclosable faults on exactly that plant. And it will not survive a facilities team with no allocated capacity, whatever the business case said.

The idea to walk away with

FDD works, and in buildings it works best in its least fashionable form: explicit engineering rules, written against known physics, producing findings a mechanical engineer can act on without a data scientist to interpret them. The fault library is well established, the recurring faults are the same everywhere, and a meaningful share of them cost nothing but engineering hours to fix. That is an unusually good return profile.

What determines whether you get that return is not the tool. It is three things: whether the trend data exists at the interval and retention the rules need, whether the output arrives as a ranked short list rather than a wall of alerts, and whether a named person with allocated hours owns closing that list through the CMMS. Get those right and a modest rule set on one building pays for itself and funds the rollout. Get them wrong and the best platform on the market is a licence nobody logs into.

Final thoughts

If you are considering FDD, the most valuable work happens before procurement and costs almost nothing. Export your points list. Take the twenty rules you would most want to run and check whether the points exist, are trended at a usable interval and are retained long enough. Establish a consumption baseline. Decide who owns the findings and how many hours a week they have. Write the CMMS integration into the requirement rather than a later phase. Then, and only then, look at platforms.

Done in that order, FDD is one of the few building technology investments where the first year genuinely pays and the results are defensible when finance asks. Done platform-first, it joins the long list of correctly functioning software that changed nothing. Continuous commissioning is a habit before it is a product, and the buildings that do it well are not the ones with the best analytics. They are the ones where someone looks at a short list every week and fixes what is on it.

Disclosure

Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.

Planning an FDD or building analytics programme?

Independent advisory on FDD data readiness, rule library scoping, fault prioritisation, CMMS integration of findings, and the measurement and verification framework to prove the savings. 22+ years across enterprise CMMS, CAFM, EAM and ERP implementations.

Book a conversation

Related reading: Building management systems: a complete guide, Smart buildings: from BMS to building analytics, Energy management systems for buildings, BMS in HVAC: controls, points and sequences, AI anomaly detection and early fault warning, AI for energy, water and carbon.

Muhammad Abbas

CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.

Work with me
MAbbaz.com
© MAbbaz.com