Key takeaways
See our roundup of RCA software for these patterns.
The pattern is simple and ubiquitous. A specific failure mode hits an asset, say, a sensor false-trigger on Line 3 packaging head. The technician resets the sensor; the line restarts; the work order closes. Three weeks later the same false-trigger hits. Same fix. Six weeks later, again.
By the time the failure has hit five or six times over a quarter, the team's mental model has accommodated it. "The packer's sensor is finicky." Operators have a workaround. Maintenance has a fast-response procedure. The cost is real, production minutes lost on each event, plus the unmeasured drag of the workaround, but no single event is bad enough to escalate.
This is the dynamic the threshold rule is designed to break. Without an explicit rule, the team's natural calibration is to handle each event individually because each event individually is manageable. The cumulative cost stays invisible.
The rule: any failure mode that recurs three times on the same asset class within a 90-day window triggers a structural investigation. Not a fix, an investigation, with a specific protocol described below.
Two occurrences is often just noise, random variation in operating conditions can produce two of the same failure without a common cause. Three within a tight window is a practical threshold: conservative enough that ordinary noise rarely trips it, aggressive enough to catch a real pattern early.
A 12-month window is too forgiving, the failure mode has been running for nine months by the time it triggers, which is past the point where intervention is cheap. 90 days catches the pattern while the operational context is still fresh and the cause is still investigable.
An asset-class trigger catches systemic causes (the same part fails the same way across multiple assets); an asset-only trigger only catches one-off configuration problems. Most high-leverage interventions are at the class level. The piece on manufacturing KPIs covers the related cross-asset trend metrics.
The word makes it sound dramatic. It is not. Escalation in this framework is a 4-hour structured investigation, not a crisis meeting. The protocol:
The reliability engineer or maintenance manager pulls every work order on the failure mode in question over the trigger window. They look at: when each event happened (time of day, day of week, position in the run), what was running on the asset, what was done to repair, what was consumed.
From the data, one or two specific hypotheses about the cause. Not a brainstorm. A short list of testable explanations, with the data pattern that would confirm or refute each. The piece on root cause analysis covers hypothesis formation in more detail.
Investigation moves from the data to the physical asset. The reliability engineer (or designated investigator) walks the asset with the technician who has been doing the repairs. The goal is to see what the data cannot show, wear patterns, cable routing, environmental factors, operator interactions.
The output is a one-page memo: the failure mode, the data pattern, the hypothesis, the proposed intervention, the named owner, and the expected resolution date. The memo goes into the CMMS attached to the asset class so the next occurrence is automatically aware of the prior investigation.
That is the entire escalation. Four hours, one investigator, one technician, one memo. Done within two weeks of the third occurrence triggering the rule.
The 3-in-90 rule will sometimes trigger on real noise, three events that look like a pattern but were coincidence. The investigation runs anyway. The output in those cases is a memo that says "no structural cause identified after investigation; continue routine fix protocol." That memo is itself valuable, it documents that the team did look, and it raises the bar for the next claim of "this is just bad luck."
Plants worry about false-positive investigations. In practice, the time cost is small (4 hours), and the cost of a missed real pattern (failures running unchallenged) is much higher, so the rule deliberately errs on the side of looking.
The rule needs to be enforced automatically or it gets forgotten. The CMMS needs to:
Plants without this automation tend to lose the discipline within a quarter, the human-driven count just fails to keep up with the volume.
The rule works in any CMMS that tracks failure modes and dates. Where a unified OEE + CMMS platform helps is that the 3-in-90 detection is a live trigger rather than a periodic report, and the resulting investigation can pull OEE-event context (when the failures happened in the production cycle) without manual reconciliation.
Fabrico is built for that workflow. The connection upstream into the preventive maintenance schedule matters because the investigation usually concludes with a PM cadence change. To see what the 3-in-90 picture would surface against your last year's data, book a demo .
Yes. Safety-critical failure modes get a 1-in-30 rule: any recurrence within 30 days triggers immediate investigation. The threshold is tighter because the cost of accumulation is qualitatively different.
The trigger is automatic. Whether the team agrees with the conclusion of the investigation is a separate matter. The investigation runs; the conclusion can be "no structural cause found"; that is a valid outcome. The rule is not "we must find something"; it is "we must look."
Single named investigator with capacity allocated explicitly. Plants that try to assign investigations to whichever engineer is available usually end up with a backlog. One person, with three to six investigations per quarter on their explicit workload, executes consistently.
The rule is per failure mode, not per asset. A packaging head can have ten different failure modes; each one tracks against its own 3-in-90 counter. This is why the failure mode catalogue matters, without consistent mode definitions, the counters drift.
Letting the trigger fire without the investigation being run. The first time the team skips an investigation because the maintenance manager judges it "not worth it," the rule starts to decay. The plant manager has to enforce that every trigger produces an investigation, even when the conclusion is "no structural cause." The discipline is what makes the rule work.