Menu
Failure Mode Catalogues: How to Build One Without It Becoming Bureaucracy

Failure Mode Catalogues: How to Build One Without It Becoming Bureaucracy

Most failure mode catalogues die from over-scoping. A bounded method: top 20 asset classes, 3-7 modes each, harvested from CMMS history, wired into.
Failure Mode Catalogues: How to Build One Without It Becoming Bureaucracy

Key takeaways

See our roundup of root cause analysis software.

  • A failure mode catalogue is the most valuable maintenance artefact most plants do not have. It is also the most-abandoned project most plants have ever started, usually killed by trying to be too complete on day one.
  • A useful catalogue does not need to be exhaustive. It needs to cover the top 20 asset classes, with three to seven failure modes each, and one paragraph per mode. That is roughly 60-140 entries, achievable in two months of evenings, not a year of meetings.
  • The single biggest mistake in building a catalogue is starting with new failure modes invented in workshops. The right approach is the opposite: harvest the modes from the last 12 months of work orders, where they already exist as messy free text.
  • The catalogue earns its keep when it is wired into the CMMS, every new work order has to pick a mode, no free-text exceptions. Without that wire, the catalogue is a binder no one opens.

Why most failure mode catalogues die

Every reliability initiative includes "build a failure mode catalogue." Very few finish.

The pattern is consistent: the team commits to comprehensiveness on day one, schedules workshops to define every conceivable failure mode for every asset class, runs out of energy by month three, and the partial catalogue sits in a SharePoint folder where no one references it.

By month six the team has moved on and the catalogue's existence is forgotten.

The pattern is not a discipline problem; it is a scoping problem. A catalogue that tries to be exhaustive is impossible to keep current. A catalogue that covers the assets that matter, at the depth they need, is achievable, and the second is what actually drives improvement.

What "useful" looks like

A useful failure mode catalogue has four properties:

  • Bounded scope. Top 20 asset classes by failure frequency, no more. The long tail can wait.
  • Bounded depth. Three to seven modes per asset class. Modes 8 and beyond do not differentiate enough to be worth the maintenance burden.
  • Bounded language. One paragraph per mode, in plain English, written so the maintenance technician on the floor can match it to what they see.
  • Wired into the CMMS. Every closed work order on a covered asset class must pick a mode from the catalogue. No free-text fallbacks.

That is the whole thing. 60 to 140 entries. A reliability engineer working alone can produce it in two months. With one collaborator from the maintenance team, faster.

How to build it: harvest before invent

The wrong way to start is a whiteboard session asking the team "what could fail on this asset?" The team will produce a list dominated by recency bias and dramatic-but-rare modes, and miss the boring high-frequency ones.

The right way is to harvest from the CMMS history. Pull the last 12 months of closed work orders on the top 20 asset classes, group them by asset class, and read the closeout text. The actual failure modes are already in there, expressed in the technicians' own language, often inconsistently, often misspelt, but accurate.

The article on the work order management system covers the data structure this harvest depends on.

Cluster the messy free-text reasons into 3-7 modes per asset class. For a packaging head, the clusters might be: jaw wear, seal contamination, drive misalignment, sensor false-trigger, pneumatic leak, programming corruption. For a drive motor: bearing failure, winding short, brush wear, overload trip, cooling failure. Each cluster gets a name and a one-paragraph description.

What this approach catches that workshops miss: the high-frequency boring modes. What it misses that workshops would have caught: rare-but-catastrophic modes. The second can be added in a small follow-up pass, but the first is where the volume of preventable failures lives.

The one-paragraph mode description

Each mode gets the same paragraph structure. Roughly 80-120 words.

  • Mode name, short, descriptive, written the way a technician would say it. "Jaw wear" not "Progressive material loss at the sealing interface."
  • What the operator or technician sees, two or three sentences of observable symptoms. "Seals start showing inconsistent texture, then leaks; the line typically runs another 4-8 hours before a full stop." This is what the catalogue is for; without it, the entries are abstract.
  • Typical cause, one sentence. Not exhaustive; the most common explanation.
  • What to do when you see it, one sentence, action-oriented. "Open work order for jaw inspection within the shift; if not addressed, expect a full stop within 24 hours."

The piece on root cause analysis goes deeper on the cause investigation; the catalogue intentionally stays shallow.

Wiring it into the CMMS

A catalogue that lives in a binder is decoration. A catalogue that lives in the CMMS, where every new work order on a covered asset must pick a mode, is the artefact that earns its keep. Three rules for the wiring:

  1. Required field, no closed work order without a selected mode. Free text is allowed as a supplement, never a substitute.
  2. An "other / not in catalogue" option, but it must trigger a review. Three "other" entries on the same asset class in a quarter mean the catalogue is missing a mode; add it.
  3. Quarterly catalogue review, counts per mode, plus the "other" entries. Modes that have not fired in 12 months get retired. Modes that fired more than expected get a closer look. The article on the preventive maintenance schedule covers how mode counts feed back into PM cadence.

The "other" trigger is the mechanism that keeps the catalogue current without anyone consciously maintaining it. It runs on the principle that the right time to add a mode is when it has fired three times, not when someone imagined it in a workshop.

What the catalogue produces

Three things, in order of value.

1. Better trend data

Once every work order has a mode, the failure-rate-per-mode analysis becomes possible. Jaw wear on packaging head P3 fires 23 times this quarter, up from 14 last quarter, that is a trend the team can act on. Without the catalogue, the trend would have been "more packaging head failures," which is too vague to act on. The article on manufacturing KPIs covers the trend metrics this enables.

2. Faster triage on the floor

A technician walking up to a stopped asset can scan three to seven modes for that asset class, match symptoms to one, and start the right repair immediately. The triage time falls from "10 minutes of diagnosis" to "60 seconds of mode-matching." For high-frequency failure modes the gain compounds across the year.

3. PM tuning that actually works

When a mode fires repeatedly, the corresponding PM task can be tightened (more frequent, more thorough, different parts). Without the catalogue, PM tuning runs on aggregate failure counts and tends to over-tighten everything; with it, the tightening matches the mode.

How Fabrico fits

The catalogue method works in any CMMS that allows a required dropdown on work-order closure. Where a unified OEE + CMMS platform helps is the harvest step: pulling 12 months of work orders with their OEE-event context is one query rather than a manual export and reconciliation.

Fabrico is built so the catalogue lives in the same database as the work orders, the OEE events and the asset hierarchy, modes can be added or retired without breaking the trend history. To see what a working catalogue looks like for your top asset classes, book a demo .

Frequently asked questions

Should we use a standard taxonomy like ISO 14224?

Reference it; do not adopt it wholesale. ISO 14224 originated in oil-and-gas reliability reporting and contains modes that do not apply to most manufacturers, while missing modes that do. Use it as a checklist when reviewing your harvested modes, it catches the modes you forgot, but write the catalogue in your team's language.

How many modes per asset class is the right number?

Three to seven. Below three, the catalogue is too coarse to differentiate causes. Above seven, modes blur together and technicians stop using them. If an asset class genuinely has more, split it into sub-classes.

What if our work-order history is too messy to harvest from?

That is the first project: clean up the closeout fields on the top 20 asset classes' work orders for the last 12 months. This is unglamorous and takes 4-6 weeks; without it, the catalogue is built on guesses. The output of the cleanup is the input to the catalogue.

Who owns the catalogue?

The reliability engineer, or in plants without that role, the maintenance manager. Single owner. Quarterly review. Changes require a one-sentence rationale in the change log.

What is the most common failure mode of the failure mode catalogue?

Letting it become a separate document instead of a CMMS dropdown. A catalogue that requires a binder lookup is used for two weeks. A catalogue that is the dropdown in the work order is used every day. Wire it into the CMMS or do not build it.

Latest from our blog

Define Your Reliability Roadmap
Validate Your Potential ROI: Book a Live Demo
Define Your Reliability Roadmap
By clicking the Accept button, you are giving your consent to the use of cookies when accessing this website and utilizing our services. To learn more about how cookies are used and managed, please refer to our Privacy Policy and Cookies Declaration