Menu
When to Escalate a Recurring Problem: The Threshold Rule

When to Escalate a Recurring Problem: The Threshold Rule

Recurring problems eat capacity slowly because every individual occurrence is manageable. The 3-in-90 rule triggers structural investigation before.
When to Escalate a Recurring Problem: The Threshold Rule

Key takeaways

See our roundup of RCA software for these patterns.

  • Recurring problems eat capacity slowly. The maintenance team treats each occurrence as a routine fix, the production team treats each occurrence as bad luck, and the underlying cause runs unchallenged for quarters. Most plants have at least one example.
  • The escalation problem is decision-shaped: at what point does "we'll handle it again this time" become "this needs a structural investigation"? Without an explicit rule, the answer is always "next time", and there is always a next time.
  • The threshold rule: any failure mode that recurs three times on the same asset class within 90 days is escalated to a structural investigation regardless of impact. The 3-in-90 rule is conservative enough that real noise does not trigger it and aggressive enough that real patterns get caught early.
  • Escalation does not mean a crisis. It means a 4-hour structured investigation with a named owner, a written hypothesis, and a measurable intervention. The point is to break the routine-fix loop before the failure mode embeds in the operational pattern.

The slow accumulation of unchallenged failures

The pattern is simple and ubiquitous. A specific failure mode hits an asset, say, a sensor false-trigger on Line 3 packaging head. The technician resets the sensor; the line restarts; the work order closes. Three weeks later the same false-trigger hits. Same fix. Six weeks later, again.

By the time the failure has hit five or six times over a quarter, the team's mental model has accommodated it. "The packer's sensor is finicky." Operators have a workaround. Maintenance has a fast-response procedure. The cost is real, production minutes lost on each event, plus the unmeasured drag of the workaround, but no single event is bad enough to escalate.

This is the dynamic the threshold rule is designed to break. Without an explicit rule, the team's natural calibration is to handle each event individually because each event individually is manageable. The cumulative cost stays invisible.

The 3-in-90 rule

The rule: any failure mode that recurs three times on the same asset class within a 90-day window triggers a structural investigation. Not a fix, an investigation, with a specific protocol described below.

Why three and not two

Two occurrences is often just noise, random variation in operating conditions can produce two of the same failure without a common cause. Three within a tight window is a practical threshold: conservative enough that ordinary noise rarely trips it, aggressive enough to catch a real pattern early.

Why 90 days and not a full year

A 12-month window is too forgiving, the failure mode has been running for nine months by the time it triggers, which is past the point where intervention is cheap. 90 days catches the pattern while the operational context is still fresh and the cause is still investigable.

Why "same asset class" and not "same asset"

An asset-class trigger catches systemic causes (the same part fails the same way across multiple assets); an asset-only trigger only catches one-off configuration problems. Most high-leverage interventions are at the class level. The piece on manufacturing KPIs covers the related cross-asset trend metrics.

What "escalation" actually means

The word makes it sound dramatic. It is not. Escalation in this framework is a 4-hour structured investigation, not a crisis meeting. The protocol:

Hour 1: Pull the data

The reliability engineer or maintenance manager pulls every work order on the failure mode in question over the trigger window. They look at: when each event happened (time of day, day of week, position in the run), what was running on the asset, what was done to repair, what was consumed.

Hour 2: Form a hypothesis

From the data, one or two specific hypotheses about the cause. Not a brainstorm. A short list of testable explanations, with the data pattern that would confirm or refute each. The piece on root cause analysis covers hypothesis formation in more detail.

Hour 3: Walk the asset

Investigation moves from the data to the physical asset. The reliability engineer (or designated investigator) walks the asset with the technician who has been doing the repairs. The goal is to see what the data cannot show, wear patterns, cable routing, environmental factors, operator interactions.

Hour 4: Document and assign

The output is a one-page memo: the failure mode, the data pattern, the hypothesis, the proposed intervention, the named owner, and the expected resolution date. The memo goes into the CMMS attached to the asset class so the next occurrence is automatically aware of the prior investigation.

That is the entire escalation. Four hours, one investigator, one technician, one memo. Done within two weeks of the third occurrence triggering the rule.

What happens if the investigation is wrong

The 3-in-90 rule will sometimes trigger on real noise, three events that look like a pattern but were coincidence. The investigation runs anyway. The output in those cases is a memo that says "no structural cause identified after investigation; continue routine fix protocol." That memo is itself valuable, it documents that the team did look, and it raises the bar for the next claim of "this is just bad luck."

Plants worry about false-positive investigations. In practice, the time cost is small (4 hours), and the cost of a missed real pattern (failures running unchallenged) is much higher, so the rule deliberately errs on the side of looking.

Where to set the rule into the CMMS

The rule needs to be enforced automatically or it gets forgotten. The CMMS needs to:

  • Track failure modes per asset class (depending on the work order management system and the failure mode catalogue).
  • Detect the 3-in-90 pattern on close-out of any work order.
  • Open a tagged "escalation" task assigned to the named investigator.
  • Surface the open escalation task in the weekly maintenance review until it is closed.

Plants without this automation tend to lose the discipline within a quarter, the human-driven count just fails to keep up with the volume.

How Fabrico fits

The rule works in any CMMS that tracks failure modes and dates. Where a unified OEE + CMMS platform helps is that the 3-in-90 detection is a live trigger rather than a periodic report, and the resulting investigation can pull OEE-event context (when the failures happened in the production cycle) without manual reconciliation.

Fabrico is built for that workflow. The connection upstream into the preventive maintenance schedule matters because the investigation usually concludes with a PM cadence change. To see what the 3-in-90 picture would surface against your last year's data, book a demo .

Frequently asked questions

Should the rule be tighter for safety-critical failures?

Yes. Safety-critical failure modes get a 1-in-30 rule: any recurrence within 30 days triggers immediate investigation. The threshold is tighter because the cost of accumulation is qualitatively different.

What if the team disagrees with the trigger?

The trigger is automatic. Whether the team agrees with the conclusion of the investigation is a separate matter. The investigation runs; the conclusion can be "no structural cause found"; that is a valid outcome. The rule is not "we must find something"; it is "we must look."

How do we keep the investigations from piling up?

Single named investigator with capacity allocated explicitly. Plants that try to assign investigations to whichever engineer is available usually end up with a backlog. One person, with three to six investigations per quarter on their explicit workload, executes consistently.

What about asset classes with many possible failure modes?

The rule is per failure mode, not per asset. A packaging head can have ten different failure modes; each one tracks against its own 3-in-90 counter. This is why the failure mode catalogue matters, without consistent mode definitions, the counters drift.

What is the most common implementation mistake?

Letting the trigger fire without the investigation being run. The first time the team skips an investigation because the maintenance manager judges it "not worth it," the rule starts to decay. The plant manager has to enforce that every trigger produces an investigation, even when the conclusion is "no structural cause." The discipline is what makes the rule work.

Latest from our blog

Define Your Reliability Roadmap
Validate Your Potential ROI: Book a Live Demo
Define Your Reliability Roadmap
By clicking the Accept button, you are giving your consent to the use of cookies when accessing this website and utilizing our services. To learn more about how cookies are used and managed, please refer to our Privacy Policy and Cookies Declaration