Menu
Equipment Failure Records: The Reliability Data Model

Equipment Failure Records: The Reliability Data Model

An equipment failure record captures the asset, failure mode, cause, downtime, and repair of each breakdown.
Equipment Failure Records: The Reliability Data Model

Key takeaways

  • An equipment failure record is the structured log of a single breakdown: the asset, the failure mode, the cause, detection, downtime start and end, repair time, and parts consumed.
  • Failure records are only useful when failure mode (the observed symptom) and cause (the underlying reason) are captured as separate, coded fields, not free text.
  • Clean records are the raw material for MTBF, MTTR, OEE availability, and FMEA. Garbage records make every downstream reliability metric unreliable.
  • The biggest data-quality risk is a guessed cause. Capturing the true cause of downtime at the moment it happens keeps the whole reliability data model honest.

What is an equipment failure record?

An equipment failure record is the structured documentation of a single failure event on a single asset. It answers what failed, how it failed, why it failed, how long it was down, what was done to fix it, and what parts were consumed. One record equals one breakdown, timestamped and coded so it can be counted, sorted, and analyzed later.

A failure record is not the same as a work order. The work order is the instruction to do the repair. The failure record is the reliability evidence the repair leaves behind.

Without accurate failure data, a CMMS is little more than a work order ticket system, and without failure modes its codes say little about how equipment actually fails. The record is what turns a closed ticket into data you can learn from.

This matters because nearly every reliability metric your plant reports is built on these records. If the records are vague, guessed, or inconsistently coded, then MTBF, MTTR, and the availability component of OEE are all built on sand.

Which fields belong in an equipment failure record?

A complete failure record captures the asset, the failure itself, the timing, the response, and the resources. Missing any one of these breaks a downstream calculation. The table below shows the core fields, what each one feeds, and the common failure of that field.

FieldWhat it capturesWhat it feedsCommon data-quality problem
Asset IDThe exact equipment, down to component levelPer-asset Pareto, criticality rankingLogged against a parent line, not the failing unit
Failure modeThe observed symptom (e.g. bearing seized, motor overheated)FMEA, failure-pattern analysisFree text instead of a coded value
CauseThe underlying reason (e.g. misalignment, contamination)Root cause, recurring-failure preventionGuessed, blank, or conflated with the mode
Detection methodHow the failure was found (operator, alarm, inspection)Detectability scoring in FMEANot recorded at all
Downtime start / endWhen the asset stopped and resumed productionOEE availability, total downtimeRounded or backfilled from memory
Time to repairActive wrench time to restore functionMTTRConfused with total downtime
Parts consumedComponents and quantities usedCost analysis, spares planningLogged after the fact, incomplete
Action takenThe corrective work performedRemedy library, knowledge reuseMerged into the cause field

A subtle but critical point: downtime and time to repair are different fields. Downtime is the full clock from stop to restart, including waiting for a technician and waiting for parts. Time to repair is the active repair work only. Storing them separately is what lets you later distinguish a slow-to-respond problem from a hard-to-fix problem.

Why must failure mode and cause be separate fields?

Because the failure mode is what a technician observes and the cause is what they infer. The mode is usually correct. The cause is often a best guess. If you merge them into one box, you can never tell a reliable observation from a hopeful assumption, and your data quality silently degrades.

Keeping them apart also enables the most valuable reliability question of all: does the same mode keep recurring from the same cause? The trap is familiar: a new bearing goes in, the true cause is never resolved, and the replacement fails the same way a few months later. Separate fields are what surface that pattern.

What is a failure code and why does taxonomy matter?

A failure code is a standardized, dropdown-selectable value that classifies a failure instead of describing it in free text. A coded taxonomy turns thousands of breakdowns into a countable dataset. Free text turns them into a pile of unsearchable notes.

The reference standard here is ISO 14224 , the international standard for the collection and exchange of reliability and maintenance data for equipment.

It defines a common reliability language: a layered equipment taxonomy plus standardized failure-data categories covering equipment data, failure data such as failure cause and consequence, and maintenance data such as maintenance action and down time.

Even outside oil and gas, where it originated, ISO 14224 is the benchmark for how to structure failure records so they pool into one statistically valid dataset ( ISO 14224:2016 ).

The practical lesson for any plant: build a controlled vocabulary of failure modes and causes per equipment class, enforce it with dropdowns, and make the failure-mode field mandatory before a corrective work order can close. That single rule does more for data quality than any analytics tool.

How do failure records feed MTBF, MTTR, and OEE?

Failure records are the input data for the headline reliability metrics. Each metric is a simple ratio, but every term in that ratio comes straight off your records, which is exactly why record quality determines metric quality.

MetricFormulaRecord fields it consumes
MTBF (Mean Time Between Failures)Total uptime / number of failuresFailure count, operating time between records
MTTR (Mean Time To Repair)Total repair time / number of repairsTime-to-repair field, failure count
OEE AvailabilityRun time / planned production timeDowntime start and end on every record

The formulas are deliberately simple. MTBF is total uptime divided by the number of failures, and MTTR is total repair time divided by the number of repairs ( LogicMonitor ). Notice that both depend on an accurate count of failures.

If small stops go unrecorded, your failure count is too low and MTBF looks artificially good. If downtime timestamps are sloppy, availability is wrong. For a deeper treatment, see our guides on MTBF and MTTR and on unplanned downtime .

How do clean records enable FMEA and reliability analysis?

A Failure Mode and Effects Analysis is only as good as the failure history it is built from. FMEA ranks risks using severity, occurrence, and detectability. Your failure records supply the real-world occurrence data (how often each mode actually happens) and the detection data (how the failure was found). Without coded records, those scores are opinions. With them, they are evidence.

Clean records also drive asset criticality ranking and let you build a data-led preventive maintenance program instead of a calendar-based guess. A recurring failure mode on a critical asset is the textbook trigger to convert reactive repair into a planned PM task.

What are the most common failure-record data-quality problems?

The most common problem is the guessed cause: a technician closes the job, picks a plausible cause from the dropdown, and moves on. The second is missing records entirely, where short stops never get logged. Together they corrupt both the cause data and the failure count, which is why so many CMMS installations never produce a usable failure analysis.

Use this checklist to audit your own failure-record discipline:

  • Mandatory failure mode on every corrective work-order close, selected from a controlled list.
  • Separate fields for failure mode, cause, and action taken, never one merged note.
  • Timestamps captured at the event, not reconstructed from memory at shift end.
  • Short stops counted, so the failure count is real (see the six big losses framework).
  • A controlled taxonomy per equipment class, ideally aligned to ISO 14224 structure.
  • Cause based on evidence, not the most convenient dropdown value.

How does Fabrico keep failure records accurate, not guessed?

Most CMMS failure records are typed in after the fact by a technician who is reconstructing what happened. Fabrico changes the source of the data. Because Fabrico connects directly to machine PLCs, the downtime start and end timestamps are captured automatically from the line, not entered by hand, so the availability data feeding OEE is precise rather than rounded.

For the hardest field of all, the cause, Fabrico can capture the stop itself on video at the moment it occurs, so the technician confirms the cause from the footage instead of reconstructing it from memory. That tackles the single biggest data-quality risk in any reliability data model: a guessed cause logged hours later.

The stop can then trigger a follow-up task on the technician's phone, with QR-enforced checklists, so the corrective action and the spare parts used are recorded against the work order as the job is done, not remembered afterward.

The result is a closed fault-to-fix loop where the failure record is a byproduct of the work, not an afterthought. As an EU-built platform with EU data residency, Fabrico keeps that reliability data inside a clear governance boundary.

If your failure records are full of guessed causes and rounded timestamps, your MTBF and OEE numbers are guesses too. See how Fabrico captures every stop and turns every breakdown into clean, analysis-ready data.

A high repeat failure rate is a signal that root causes are not being fixed.

Worked example: three records, two metrics

Take a filler with 720 planned production hours in a month and three breakdowns. Downtime was 2.0, 1.5, and 4.5 hours, and active repair time was 1.2, 0.8, and 3.0 hours. Uptime is 720 minus 8, or 712 hours, so MTBF is 712 divided by 3, about 237 hours, and availability is 712 divided by 720, about 98.9 percent. MTTR is 5.0 divided by 3, about 1.7 hours, yet the average stop lasted about 2.7 hours. The one hour gap per stop is time spent waiting for a technician or for parts, and you can only see it because downtime and repair time are stored as separate fields.

Frequently asked questions

What is the difference between an equipment failure record and a work order?

A work order is the instruction to perform a repair. An equipment failure record is the structured reliability evidence that repair leaves behind: the asset, failure mode, cause, downtime, repair time, and parts. One records what to do, the other records what happened and why, so it can be counted and analyzed later.

What fields should an equipment failure record contain?

At minimum: asset ID, failure mode (the observed symptom), cause (the underlying reason), detection method, downtime start and end, time to repair, parts consumed, and action taken. Failure mode and cause must be separate, coded fields, and downtime must be stored separately from active repair time.

Why must failure mode and failure cause be separate fields?

Because the failure mode is what a technician observes and is usually correct, while the cause is what they infer and is often a best guess. Keeping them separate lets you tell reliable observations from assumptions and surface whether the same mode keeps recurring from the same unresolved cause.

How do failure records feed MTBF and MTTR?

MTBF equals total uptime divided by the number of failures, and MTTR equals total repair time divided by the number of repairs. Both pull directly from failure records: the failure count, the operating time between failures, and the time-to-repair field. Inaccurate or missing records make both metrics wrong.

What is ISO 14224 and why does it matter for failure records?

ISO 14224 is the international standard for collecting and exchanging reliability and maintenance data for equipment. It defines a common equipment taxonomy and standardized failure-data categories, including failure cause, consequence, and down time, so records pool into one statistically valid dataset. It is the benchmark for structuring failure records consistently.

What is the most common data-quality problem in failure records?

The most common problem is a guessed cause: the technician picks a plausible value from the dropdown after the fact instead of capturing the true cause. Combined with unrecorded short stops, this corrupts both cause data and the failure count, which is why many CMMS installations never produce a usable failure analysis.

Latest from our blog

Define Your Reliability Roadmap
Validate Your Potential ROI: Book a Live Demo
Define Your Reliability Roadmap
By clicking the Accept button, you are giving your consent to the use of cookies when accessing this website and utilizing our services. To learn more about how cookies are used and managed, please refer to our Privacy Policy and Cookies Declaration