mail@mabbaz.com Abu Dhabi, UAE

Failure Analysis · Reliability · Maintenance Engineering

Failure Analysis: Methods, Process and Examples

Root cause analysis interrogates the system and the organisation. Failure analysis interrogates the part. This guide is about the physical examination of a component that has actually failed: how to preserve the evidence before well-meaning people destroy it, how to read fracture faces and worn surfaces at a practitioner's level, how to separate the failure mode from the mechanism from the cause, and when the honest answer is to send it to a laboratory.

Muhammad Abbas September 27, 2026 ~21 min read

The most valuable object in any investigation is the broken part itself, and in most maintenance organisations it is in a skip before lunch. In review after review the discussion runs for two hours on process, blame, spares policy and operator behaviour, and then somebody asks the only question that would have settled it: where is the failed component? Nobody knows. It was cleaned, it was replaced, the debris was swept up, the oil was drained to the waste tank, and the only photograph is of the new part fitted. That is the failure of failure analysis, and it happens before the analysis starts.

The message up front: failure analysis is an evidence discipline. The failed component is a physical record of how it died, written in fracture faces, wear patterns, discolouration, debris and deposits, and that record is fragile. Most of it is destroyed in the first hour by people trying to be helpful. Get the preservation right and even a modest visual examination will tell you something useful. Get it wrong and no amount of workshop facilitation later will recover what the part could have told you.

1. What failure analysis actually is

Failure analysis is the examination of a failed item to establish how it failed: what physically happened to the material, in what sequence, and under what loading or environmental condition. It is an engineering activity performed on an object. Its raw material is the component, the mating parts, the debris, the lubricant, the operating record at the moment of failure and what the operator saw, heard or smelled.

That is a narrower activity than most people mean when they say "we did a failure analysis". In many organisations the phrase has drifted to mean any structured discussion after a breakdown. It is worth reclaiming the narrow meaning, because the narrow activity is the one that is rarely done, and it is the one that produces findings the rest of the investigation cannot argue with. A cause-and-effect workshop can produce three plausible stories. A fracture face usually supports one.

This is also not condition monitoring. Condition monitoring watches a running asset for signs that a failure is developing, which is the territory of the P-F curve and of vibration, thermography and oil analysis. Failure analysis happens after the event, on the wreckage. The two feed each other: a physical finding tells you what signature you should have been watching for, and a monitoring history tells the examiner how long the damage was developing.

2. How failure analysis differs from root cause analysis

These are complementary activities with different subjects, and confusing them is the reason so many investigations stall.

  • Root cause analysis reasons about the system. It asks why the conditions that destroyed the part were allowed to exist: why the alignment was never checked, why the lubrication route was dropped, why the wrong part was in the storeroom, why the alarm was disabled. It reaches into procedures, competence, design decisions and management systems. There is a genuine international standard in this space, IEC 62740:2015 "Root cause analysis (RCA)", which sets out RCA principles and process steps and describes a set of recognised techniques, including the "Why" method and the Ishikawa or fishbone diagram, and which explicitly excludes assigning blame.
  • Failure analysis interrogates the part. It asks what physically happened to this piece of metal, elastomer, winding or bearing surface. Its output is a mechanism and a loading condition, not an organisational finding.

The relationship is sequential and one-directional in a useful way: failure analysis constrains RCA. If the examination shows a bearing that was starved of lubricant, the RCA has a defined problem to explain and a whole family of speculative causes is eliminated before the meeting starts. If it shows a fatigue fracture originating at a machining mark, the conversation moves to manufacture and specification rather than operator error. Skipping the physical step is what lets an RCA wander for weeks. For the investigation process itself, the technique landscape and how to run the sessions, see the root cause analysis methods guide. This article deliberately stays on the component.

3. Mode, mechanism, cause: three words people use interchangeably

Practitioners lose more time to this confusion than to any technical difficulty, so it is worth being precise.

  • Failure mode is the way in which the item stopped doing its job, stated functionally. The pump does not deliver flow. The shaft does not transmit torque. The seal does not retain fluid. It is observable from outside the component and it is what the operator reports. The full taxonomy of mode types is covered in failure modes: common types and how to analyse them.
  • Failure mechanism is the physical, chemical or metallurgical process that produced that mode. Fatigue crack propagation. Adhesive wear. Pitting corrosion. Cavitation erosion. Electrical discharge damage. Thermal degradation of an elastomer. This is what the examination of the part establishes, and it is the specific contribution of failure analysis.
  • Cause is why that mechanism was active. Misalignment imposing a bending load. A lubricant that was the wrong viscosity. Suction conditions that allowed vapour bubbles to form. A shaft current path with no bearing insulation or earthing brush.

A findings sentence that contains all three is useful: the coupling failed to transmit torque (mode) through reversed bending fatigue initiating at the keyway (mechanism) because the driven machine was misaligned beyond tolerance after the last overhaul (cause). A sentence that stops at the mechanism is where most reports end, and it is the single most common reason a failure recurs. "Bearing failed due to fatigue spalling" is a true statement that changes nothing.

This vocabulary is also where your maintenance history is either usable or not. If the fields in your system collapse mode, mechanism and cause into one free-text box, you will never aggregate anything. The structure that fixes it is set out in failure codes: problem, cause, action, and the international reference for reliability and maintenance data collection and exchange is ISO 14224:2016, whose equipment and failure taxonomy is worth reading even if you are nowhere near oil and gas. One correction worth making because it is repeated constantly: OREDA is a proprietary members-only database, not a standard. ISO 14224 is the standard that grew out of that work.

4. Evidence preservation: the hour that decides everything

This is the section that earns the article, because it is the part that is cheap, requires no specialist, and is routinely skipped.

A failed component is evidence. Everything on it is information: the position of the fracture, the orientation of the wear scar, the colour of the deposit, the distribution of the debris, the smell of the oil, the witness marks where something contacted something it should not have. Nearly all of it can be destroyed in minutes by a competent, conscientious technician doing exactly what they have always been trained to do, which is to get the asset back in service and leave the area clean.

The destructive acts are all well-intentioned. Wire-brushing the part so the supervisor can see it properly removes the corrosion products and the oxide colours that would have told you the operating temperature. Solvent-washing a bearing flushes away the wear debris whose size and shape identify the mechanism. Sweeping up the fragments loses the piece that carried the crack origin. Fitting the two halves of a fracture back together to show someone rubs the most informative surface in the whole investigation. Draining the lubricant to the waste tank throws away a sample that a laboratory could have read directly. Reassembling before anyone photographs anything loses the relative position of every part.

What follows is the practical list I would put on a laminated card in every workshop and store. It costs a few bags, a marker and ten minutes.

Do Do not
Photograph in situ, before anything is touched, wide shot then close, with a scale object in frame. Do not remove, lift or rotate the component before the in-situ photographs exist.
Record position and orientation: which way up, which end drove, rotation direction, clock position of the damage. Do not rely on memory or on "we will know from the drawing". Orientation is routinely lost and rarely recoverable.
Leave the part exactly as found: deposits, discolouration, corrosion products and lubricant film intact. Do not clean, brush, blast, solvent-wash or polish anything. This is the single most damaging habit.
Bag and label debris, fragments and particles separately, noting where each was found. Do not sweep up, mix debris from different locations, or discard "small bits".
Take and retain a lubricant or process fluid sample, and keep the used filter element. Do not drain to waste, top up, or flush the system before sampling.
Retain the mating and adjacent parts: the housing, the shaft, the other half of the coupling, the seal, the fasteners. Do not scrap, return under warranty or send for repair until the examination scope is agreed.
Protect fracture faces: support them, keep them dry, wrap loosely in clean material, keep them apart. Do not fit the fracture halves back together, touch the fracture surface, or let it rust in a damp store.
Capture the human account while it is fresh: what the operator heard, saw, smelled or felt, and when. Do not leave the operator interview until the following week, and do not let it become an accusation.
Pull the operating record around the event: trends, alarms, loads, starts, process conditions. Do not assume the historian keeps high-resolution data indefinitely. Extract and save it that day.
Give the part a unique reference, log who holds it, and keep a documented chain of custody. Do not leave it on a bench. Unlabelled failed parts become scrap within days.
The test that makes this stick

Write the preservation rule into the work order type for any failure above your criticality threshold, so the technician is prompted at the point of work rather than lectured afterwards. A quarantine shelf with labelled bins and a sign-out log does more for your failure analysis capability than any training course. If the part has to be retained, the corrective work order should not be closable until a retention reference is recorded.

5. The examination sequence

The order matters, because each step is capable of destroying the evidence the next step needs. Work from the least invasive outward.

  • Step 1: background, before you look. Service history, duty, time in service, what was done at the last intervention, whether this is a repeat, what changed recently. Examining a part with no context invites you to see what you expect.
  • Step 2: visual examination as received. Unaided eye and good raking light, then a hand lens. Photograph every stage. Note the location and orientation of damage relative to load direction and rotation. Look for what is absent as well as what is present: an unworn area on a supposedly loaded surface is a finding.
  • Step 3: low-magnification optical examination. A stereo microscope or even a good macro lens transforms fracture and wear reading. This is the highest-value piece of equipment a maintenance engineering function can own and one of the cheapest.
  • Step 4: dimensional and fit checks. Clearances, roundness, interference fits, runout, thread condition, the geometry of the seat. Many failures are a fit problem wearing a mechanism costume.
  • Step 5: cleaning, but only now and only deliberately. If cleaning is necessary to see a surface, decide what you are prepared to lose, photograph and sample the deposits first, and use the gentlest method that works.
  • Step 6: decide whether specialist analysis is warranted. That decision is a judgement about consequence and recurrence, covered below. It is a decision point, not a default.

Most organisations can do steps one to four competently with a camera, a lens, a scale, measuring equipment they already own and someone patient. That is not a trivial capability. A large share of repeat failures in general facilities and utility plant are diagnosable at that level, because the mechanisms involved are the common ones and the evidence is gross rather than microscopic.

6. Reading a fracture: what the broken face suggests

A fracture surface is a record of how the crack travelled, and that is why it must not be rubbed, mated or cleaned. At a practitioner's reading level, described in broad terms rather than as a metallurgical determination:

  • Fatigue tends to present as a relatively flat, comparatively smooth region with a recognisable origin, often at a stress concentration such as a keyway, fillet, thread root, corrosion pit or machining mark, with progression markings that curve away from that origin, and a final region that looks different because it broke in one go once the remaining section could no longer carry the load. The general appearance is of a crack that grew over time and then a sudden finish. Fatigue implies cyclic loading, and the useful questions are about the source of the cycling and about what created the initiation site.
  • Ductile overload tends to show visible deformation: necking, bending, shear lips, a fibrous or dull torn appearance. The material stretched before it parted. That points to a load or stress the part was never expected to see, or to a section that was reduced by earlier wear or corrosion.
  • Brittle overload tends to be flat with little or no deformation, often bright or crystalline in appearance, sometimes with markings that appear to radiate from an origin, and the pieces will usually fit back together geometrically, which is exactly why people are tempted to do it. Brittle behaviour raises questions about low temperature, material condition, heat treatment, hydrogen, or a design carrying a sharp notch.
  • Torsional failure on a shaft typically presents at an angle to the axis, or with helical features, rather than square across it. The geometry of a fracture relative to the axis is often the fastest orientation clue available.
  • Environmentally assisted cracking tends to produce branching cracks with little deformation, frequently associated with a corrosive medium and a sustained stress. This family is genuinely difficult to distinguish visually from other flat fractures and is one of the strongest triggers for laboratory work.

Two habits improve fracture reading more than any amount of reading about it. First, always establish the origin before speculating about the mechanism. Second, always ask what the load direction was, because the same appearance means different things on a part in bending, torsion or tension.

7. Reading worn and damaged surfaces

Most maintenance failures are not fractures. They are surfaces that degraded until the item stopped functioning, and surfaces carry their own vocabulary.

  • Abrasive wear commonly appears as scratching, scoring or grooving aligned with the direction of motion, produced by hard particles or a hard rough surface. It points at contamination, filtration, seal integrity and housekeeping.
  • Adhesive wear, including scuffing, galling and seizure, appears as material transfer, smearing, torn or welded-looking patches and often heat discolouration. It points at loss of lubricant film: wrong lubricant, insufficient supply, excessive load or speed, or inadequate clearance.
  • Fretting appears at nominally clamped or stationary interfaces subject to small relative movement, often as reddish-brown fine oxide debris and localised surface damage. It points at a joint that is not doing its job: a loose fit, insufficient preload, vibration through a static interface. Fretting damage is also a classic fatigue initiation site, which is how it ends up causing a fracture somewhere else.
  • Rolling contact fatigue and spalling in bearings and gears appears as pitted or flaked areas on the working surface. The distribution is the information: damage confined to one narrow band, or to one side, or at a regular spacing, usually tells you more than the damage itself.
  • Cavitation damage appears as localised spongy, honeycombed or deeply pitted areas in pump and valve internals, typically where pressure recovery occurs. It points at suction conditions, net positive suction head, throttling, entrained air or operating well away from the design duty point.
  • Erosion appears as directional loss of material where a flow of fluid or particles impinges, often with polished or wave-like surfaces and thinning at bends, nozzles and downstream of restrictions. It points at velocity, solids content or a flow path the design did not anticipate.
  • Corrosion covers a large family with different appearances: broad general attack, discrete pits, damage concentrated in crevices and under deposits, damage at the junction of dissimilar metals, or attack that follows the grain structure. Identifying the family matters because the remedies are entirely different, and pitting in particular both hides its depth and creates fatigue initiation sites.
  • Electrical discharge damage in bearings and on shafts appears as pitting, frosted or matt zones, and in some cases a regular fluted or washboard pattern on raceways. It points at shaft voltage and a current path through the bearing, a familiar issue on inverter-driven motors, and the remedies are insulation, earthing and filtering rather than anything to do with the bearing itself.

The discipline is the same in every case: describe the appearance, map where it is and is not, relate it to the direction of motion or flow, and only then propose a mechanism.

8. A working table of damage appearances

Use this as a prompt for the field, not as an identification key. The caution column is the important one.

Appearance Mechanism it suggests What to check next Caution
Flat region with an identifiable origin and progression markings, plus a final fast-break zone Fatigue under cyclic loading Source of cycling, alignment, balance, resonance, and what created the initiation site Several flat fractures look alike; a pit or corrosion feature at the origin changes the cause entirely
Necking, bending, shear lips, torn fibrous surface Ductile overload Actual load at failure, jamming, foreign object, prior section loss from wear or corrosion Overload is often a consequence of something else failing first, not the primary event
Flat, bright or crystalline, little deformation, pieces fit together Brittle fracture Temperature at failure, material and heat treatment records, notches and sharp corners, impact history Do not mate the halves to demonstrate the fit; material conclusions need laboratory work
Directional scratches and grooves along the line of motion Abrasive wear from hard particles Filtration, breather and seal condition, lubricant cleanliness, recent works generating dust Fine abrasion can be mistaken for machining marks on a surface nobody photographed when new
Smearing, material transfer, welded or torn patches, heat colours Adhesive wear, scuffing or seizure Lubricant grade and supply, load and speed, clearances, cooling, start-up conditions Heat discolouration is easily destroyed by cleaning, so record it immediately
Fine reddish-brown debris and localised damage at a clamped or static interface Fretting Preload and torque, fit condition, vibration transmitted through the joint Easily mistaken for ordinary rust; its real significance is often as a fatigue initiation site
Flaked or pitted patches on a bearing raceway or gear flank Rolling contact fatigue Where the damage sits relative to the load zone, fit, alignment, lubrication, contamination The pattern and location carry the diagnosis; the spalling itself is the end state of several different causes
Spongy, honeycombed or deeply eaten zones in pump or valve internals Cavitation Suction conditions, available NPSH, throttling and control valve position, operating point versus design duty Can resemble aggressive corrosion pitting; location relative to pressure recovery is the differentiator
Directional thinning, polished or wavy surfaces at bends and downstream of restrictions Erosion by fluid or entrained solids Velocity, solids loading, flow path geometry, upstream changes Erosion and corrosion frequently act together, so treating only one leaves the failure live
Discrete pits, crevice or under-deposit attack, damage at dissimilar metal junctions One of the corrosion families Fluid chemistry, water ingress, coating and insulation condition, material pairing, stagnant zones Pit depth is not visible from the surface, and the family matters more than the label "corrosion"
Frosted, matt or regularly fluted raceways, fine pitting on a shaft journal Electrical discharge damage Shaft voltage, earthing and brush condition, bearing insulation, drive and cabling arrangement Routinely misdiagnosed as a mechanical bearing defect, so the bearing is replaced and the failure returns
Branching cracks with little deformation in a corrosive service Environmentally assisted cracking Sustained stress, residual stress from fabrication, environment chemistry, temperature This family should go to a laboratory. Visual identification here is unreliable and the consequences are usually serious

9. The honest limits of visual identification

Where this gets people into trouble

Visual identification of failure mechanisms is genuinely difficult and frequently ambiguous. Several mechanisms look alike, damage from the final moments routinely obscures the damage that started it, secondary damage is often more dramatic than the primary event, and a clean-looking surface can hide sub-surface conditions entirely. Confident misattribution is common and is worse than an honest "not established", because it closes the investigation and the failure comes back. Where a failure has safety, environmental, contractual or legal consequence, the examination belongs with a qualified failure analyst or an accredited laboratory rather than a maintenance team's best guess, and the part should be preserved and handed over rather than examined enthusiastically in the workshop.

Within those limits, a maintenance engineer reading a component carefully is far more useful than no examination at all. The correct posture is confident description and cautious attribution. Record precisely what you observe, propose the mechanism as a hypothesis, state what would confirm or refute it, and let the strength of the conclusion follow the strength of the evidence. A report that says "consistent with fatigue initiating at the keyway, origin confirmed, cyclic source not yet established" is far more valuable than one that says "fatigue failure" with nothing behind it.

10. When to escalate to a laboratory, and what exists there

The escalation decision is driven by two things: consequence and recurrence. Escalate when the failure carried or could have carried safety or environmental consequence, when it has now happened more than once and the workshop-level explanation has already failed to stop it, when there is a warranty, insurance or contractual dispute in prospect, when the item is expensive or long-lead and getting the answer wrong means buying another one, or when the visual examination produced two competing mechanisms that cannot be separated by eye. Do not escalate a single low-consequence failure with an obvious explanation. That is how laboratory budgets get spent on the wrong things and then withdrawn.

What a laboratory can do, at orientation depth, so you can have a sensible conversation about scope:

  • Metallurgical sectioning and metallography. Cutting, mounting, polishing and etching a section to examine the internal structure, the condition of any surface treatment, and how a crack relates to the material. Note that sectioning is destructive and irreversible, which is why it comes after everything else.
  • Hardness testing. Establishes whether the material condition at the failure location is consistent with what it should be, and reveals local changes from overheating or working.
  • Microscopy. Optical microscopy for structure, and electron microscopy for fracture-surface features at a scale where mechanisms become distinguishable rather than merely plausible. This is the step that converts a visual hypothesis into a determination.
  • Spectroscopy and chemical analysis. Confirms material composition against specification and identifies deposits, corrosion products and contaminants. Frequently the fastest route to a supply-chain or counterfeit-part finding.
  • Oil and debris analysis. Wear particle counting and characterisation, contamination and water content, additive and degradation chemistry, and examination of filter and magnetic plug debris. This is the one most organisations can access routinely and most under-use. Worth knowing that personnel competence here is addressed by ISO 18436-4:2014 for field lubricant analysis and ISO 18436-5:2012 for laboratory lubricant analysis, which certify people rather than organisations or equipment, and that diagnostics as a discipline sits under ISO 13379-1:2025.

Two honest points about in-house capability. First, most organisations cannot do this work properly and should not try: the equipment, the sample preparation skill and the interpretation experience are a specialism, and a half-equipped attempt usually destroys the evidence a real laboratory needed. Second, the useful in-house investment is not a laboratory. It is a stereo microscope, a decent camera, a quarantine shelf, a preservation procedure and a relationship with an external laboratory established before you need it, so that escalation is a phone call rather than a procurement exercise.

11. Three worked examples

These are my own illustrative hypotheticals, constructed to show the reasoning. They are not client cases and the details are invented.

Example A: the motor bearing that kept coming back. A hypothetical inverter-driven pump motor fails its drive-end bearing three times in a year. Each time the bearing is replaced and the report reads "bearing failure, replaced". On the third occasion the bearing is quarantined instead of scrapped. Low-magnification examination shows a frosted raceway with a regular fluted pattern rather than the load-zone spalling you would expect from a fit or lubrication problem. Mode: motor unavailable. Mechanism: electrical discharge damage. Cause hypothesis: a shaft current path with no insulated bearing or earthing arrangement on a drive that was retrofitted. The corrective action is electrical, and no quantity of better bearings would ever have fixed it. The learning point is that the first two failures generated work orders and the third generated an answer, and the only difference was that somebody kept the part.

Example B: the impeller nobody photographed. A hypothetical chilled water pump loses head. The impeller is replaced and the old one is cleaned up for the store as a possible spare. Later the same pump degrades again. The retained impeller now shows nothing useful because it was cleaned, and nobody recorded where on the impeller the damage sat or what the suction pressure trend looked like. Had the damage been photographed in place, the spongy localised appearance near the eye would have pointed at cavitation, and the investigation would have gone to suction conditions and the operating point rather than to pump quality. Cleaning the part cost more than the impeller did.

Example C: the coupling bolt and the load direction. A hypothetical fan coupling bolt fails. The fracture face is flat over most of its area, with a discernible origin at a thread root and a different-looking final zone. The bolts either side show fine reddish-brown debris at their seating faces. Read together, the picture is fretting at an inadequately preloaded joint, cyclic loading through a flange that was moving relative to itself, and a fatigue fracture initiating where the stress concentrated. Mode: drive not transmitted. Mechanism: fatigue with a fretting initiation site. Cause hypothesis: torque procedure and joint design, not bolt quality. Note the key move: the finding came from examining the parts that had not broken.

12. Turning a physical finding into a change

A mechanism is not a deliverable. The value lands when the finding changes something durable, and there are only a handful of places for it to go.

  • A maintenance task change. A new or revised inspection, a lubrication change, a torque check, a filtration or breather change, a different interval. This is a very common landing place and the easiest to implement.
  • A condition-monitoring change. The mechanism tells you which technique would have caught it, and how early. That is how a failure improves your monitoring programme rather than just your spares consumption.
  • A design or specification change. A different material, a removed stress concentration, an insulated bearing, a coating, added redundancy, a resized line. Slower and more expensive, and sometimes the only thing that works.
  • An operating change. Keep the asset inside the duty envelope where the mechanism is not active. Often the cheapest fix available and the most likely to be overlooked.
  • A procurement or storeroom change. Correct the part specification, the supplier, the storage condition or the substitution rule that put the wrong item in service.
  • A recorded, coded history entry. Mode, mechanism and cause captured in structured fields so the next occurrence is visible as a pattern rather than as a surprise.

Two framings help decide which lever to pull. IEC 60812:2018, the third edition of the FMEA and FMECA standard, whose title changed at that edition to "Failure modes and effects analysis (FMEA and FMECA)", provides the forward-looking structure for asking what else this mechanism could affect; see the FMEA guide for how that is run, and from RUL to FMEA for the quantitative side. And task selection logic, the question of which of the levers above is actually appropriate for a given consequence, is the territory of reliability centred maintenance and of the wider discipline described in what reliability engineering is.

Software matters here only in a modest, generic way. Whatever maintenance system you run should be able to hold the structured mode, mechanism and cause fields, attach the photographs to the asset record rather than to somebody's phone, and let you retrieve every previous failure on that asset class in one query. If it cannot do those three things, your failure analysis findings will keep evaporating regardless of how well the examination was done.

13. The failure modes of failure analysis itself

The practice fails in predictable ways, and recognising your own pattern is usually enough to fix it.

  • No evidence retained. The dominant failure mode by a wide margin. Nothing downstream can compensate for it.
  • Stopping at the mechanism. "Fatigue failure" written in the report and nothing asked about why the part was being cycled. Technically correct, operationally useless.
  • Attributing too confidently. A mechanism asserted from a glance, the investigation closed, and the same failure back in six months.
  • Mistaking secondary damage for the primary event. The most spectacular damage is usually the last thing that happened, not the first.
  • Examining only the part that broke. The mating parts, the fasteners and the debris frequently carry the finding, as in example C above.
  • A report nobody reads. A careful examination written up as a PDF attached to a closed work order, never converted into a task, a specification or a monitoring change.
  • No feedback into the history. The finding never reaches structured fields, so the pattern stays invisible and the next analyst starts from zero.
  • The same failure recurring. The only real measure of whether any of this worked. If the failure comes back, the analysis was incomplete regardless of how good the report looked.

The idea to walk away with

Failure analysis is the part of investigation that interrogates the object rather than the organisation, and it is the part that constrains everything else. Its practical bottleneck is not skill or equipment. It is whether the evidence survived the first hour. A maintenance function that consistently photographs in situ, does not clean the part, bags the debris, keeps the oil, retains the mating components and quarantines the item with a label has already acquired most of the capability, and can then read fracture and wear evidence at a useful level with a lens and a microscope, escalating to a laboratory when consequence or recurrence justifies it.

Everything after that is discipline about language and follow-through: separate the mode from the mechanism from the cause, attribute no more confidently than the evidence allows, and make sure each finding lands as a task, a specification, an operating limit or a monitoring change. A mechanism identified and not acted on is a failure that is simply waiting.

Final thoughts

If you take one operational action from this article, make it the quarantine shelf and the preservation prompt on high-consequence corrective work orders. It is cheap, it needs no specialist, and it converts your failed components from waste into the most reliable evidence you will ever have about how your plant actually breaks.

The standards referenced here are voluntary and paywalled unless otherwise noted, and none of them is law by itself. IEC 62740:2015 covers root cause analysis and IEC 60812:2018 covers FMEA and FMECA, both available from the IEC . ISO 14224:2016 for reliability and maintenance data collection, ISO 13379-1:2025 for diagnostics and ISO 18436-4:2014 and ISO 18436-5:2012 for lubricant analysis personnel certification are available from ISO . Check the current published editions before you cite anything in a specification, and remember that where a failure carries safety or legal consequence, the examination belongs with a qualified specialist.

Disclosure

Alongside advisory work I also build a CMMS and CAFM platform, so I have a commercial interest in this category. Nothing above is a recommendation for it, and no vendor named here has paid for inclusion or had any editorial input. Weigh the analysis accordingly.

Same failure coming back?

Independent advisory on evidence preservation procedures, failure coding structure, and getting physical findings to land as durable maintenance and design changes. 22+ years across utilities, oil and gas, manufacturing, government and facility operations.

Book a conversation

Related reading: Root cause analysis methods, Failure modes: common types, FMEA guide, Failure codes: problem, cause, action, Condition monitoring techniques, The P-F curve explained, What is reliability engineering.

Muhammad Abbas

CMMS / CAFM Manager & Independent Advisor · 22+ years across enterprise CMMS, EAM, CAFM and ERP implementations in utilities, oil and gas, manufacturing, government and facility operations.

Work with me
MAbbaz.com
© MAbbaz.com