Key takeaways
See our roundup of root cause analysis software.
Every reliability initiative includes "build a failure mode catalogue." Very few finish.
The pattern is consistent: the team commits to comprehensiveness on day one, schedules workshops to define every conceivable failure mode for every asset class, runs out of energy by month three, and the partial catalogue sits in a SharePoint folder where no one references it.
By month six the team has moved on and the catalogue's existence is forgotten.
The pattern is not a discipline problem; it is a scoping problem. A catalogue that tries to be exhaustive is impossible to keep current. A catalogue that covers the assets that matter, at the depth they need, is achievable, and the second is what actually drives improvement.
A useful failure mode catalogue has four properties:
That is the whole thing. 60 to 140 entries. A reliability engineer working alone can produce it in two months. With one collaborator from the maintenance team, faster.
The wrong way to start is a whiteboard session asking the team "what could fail on this asset?" The team will produce a list dominated by recency bias and dramatic-but-rare modes, and miss the boring high-frequency ones.
The right way is to harvest from the CMMS history. Pull the last 12 months of closed work orders on the top 20 asset classes, group them by asset class, and read the closeout text. The actual failure modes are already in there, expressed in the technicians' own language, often inconsistently, often misspelt, but accurate.
The article on the work order management system covers the data structure this harvest depends on.
Cluster the messy free-text reasons into 3-7 modes per asset class. For a packaging head, the clusters might be: jaw wear, seal contamination, drive misalignment, sensor false-trigger, pneumatic leak, programming corruption. For a drive motor: bearing failure, winding short, brush wear, overload trip, cooling failure. Each cluster gets a name and a one-paragraph description.
What this approach catches that workshops miss: the high-frequency boring modes. What it misses that workshops would have caught: rare-but-catastrophic modes. The second can be added in a small follow-up pass, but the first is where the volume of preventable failures lives.
Each mode gets the same paragraph structure. Roughly 80-120 words.
The piece on root cause analysis goes deeper on the cause investigation; the catalogue intentionally stays shallow.
A catalogue that lives in a binder is decoration. A catalogue that lives in the CMMS, where every new work order on a covered asset must pick a mode, is the artefact that earns its keep. Three rules for the wiring:
The "other" trigger is the mechanism that keeps the catalogue current without anyone consciously maintaining it. It runs on the principle that the right time to add a mode is when it has fired three times, not when someone imagined it in a workshop.
Three things, in order of value.
Once every work order has a mode, the failure-rate-per-mode analysis becomes possible. Jaw wear on packaging head P3 fires 23 times this quarter, up from 14 last quarter, that is a trend the team can act on. Without the catalogue, the trend would have been "more packaging head failures," which is too vague to act on. The article on manufacturing KPIs covers the trend metrics this enables.
A technician walking up to a stopped asset can scan three to seven modes for that asset class, match symptoms to one, and start the right repair immediately. The triage time falls from "10 minutes of diagnosis" to "60 seconds of mode-matching." For high-frequency failure modes the gain compounds across the year.
When a mode fires repeatedly, the corresponding PM task can be tightened (more frequent, more thorough, different parts). Without the catalogue, PM tuning runs on aggregate failure counts and tends to over-tighten everything; with it, the tightening matches the mode.
The catalogue method works in any CMMS that allows a required dropdown on work-order closure. Where a unified OEE + CMMS platform helps is the harvest step: pulling 12 months of work orders with their OEE-event context is one query rather than a manual export and reconciliation.
Fabrico is built so the catalogue lives in the same database as the work orders, the OEE events and the asset hierarchy, modes can be added or retired without breaking the trend history. To see what a working catalogue looks like for your top asset classes, book a demo .
Reference it; do not adopt it wholesale. ISO 14224 originated in oil-and-gas reliability reporting and contains modes that do not apply to most manufacturers, while missing modes that do. Use it as a checklist when reviewing your harvested modes, it catches the modes you forgot, but write the catalogue in your team's language.
Three to seven. Below three, the catalogue is too coarse to differentiate causes. Above seven, modes blur together and technicians stop using them. If an asset class genuinely has more, split it into sub-classes.
That is the first project: clean up the closeout fields on the top 20 asset classes' work orders for the last 12 months. This is unglamorous and takes 4-6 weeks; without it, the catalogue is built on guesses. The output of the cleanup is the input to the catalogue.
The reliability engineer, or in plants without that role, the maintenance manager. Single owner. Quarterly review. Changes require a one-sentence rationale in the change log.
Letting it become a separate document instead of a CMMS dropdown. A catalogue that requires a binder lookup is used for two weeks. A catalogue that is the dropdown in the work order is used every day. Wire it into the CMMS or do not build it.