Menu
Mean Time to Detect: The Missing M in MTTR/MTBF

Mean Time to Detect: The Missing M in MTTR/MTBF

MTBF and MTTR are tracked everywhere. MTTD, mean time to detect, is the metric that determines whether MTTR is even measurable accurately.
Mean Time to Detect: The Missing M in MTTR/MTBF

Key takeaways

See our roundup of condition monitoring software that reduces this metric.

  • MTBF and MTTR are the maintenance metrics everyone tracks. MTTD, mean time to detect, is the one that determines whether MTTR is even measurable accurately. Most plants do not measure MTTD because they assume detection is instantaneous. It almost never is.
  • The detection gap is the time between the moment a failure or degradation actually begins and the moment the maintenance team is aware of it. In plants without explicit monitoring, this gap can be substantial, often far longer than the repair time itself, on the most common failure modes.
  • Cutting MTTD in half typically produces more avoided downtime than cutting MTTR in half on the same asset class. The same intervention is also cheaper, most MTTD gains come from sensor configuration and alert routing, not from physical changes to the asset.
  • The fix is to instrument detection at the leading-indicator level, not just at the stop event. A vibration trend, a current draw creep, a cluster of micro-stops, all happen before the asset fails outright, and each can move detection earlier by hours.

The hidden cost of slow detection

In most plants the failure timeline is told from the maintenance team's perspective. "The line went down at 14:22; the technician was on site by 14:28; the asset was running again by 14:54." Total downtime 32 minutes, MTTR 26 minutes. Clean numbers.

The maintenance manager's view is correct from the moment detection happened. What it misses is the time before that. The asset did not start failing at 14:22; it started failing at 13:15, when a sensor reading began trending out of normal range.

Nobody saw the trend, the asset kept producing degraded output for 67 minutes, then it stopped. The 32 minutes of "downtime" the team measures is the easy part. The 67 minutes of degraded production that preceded it are the part nobody counts.

This pattern is consistent enough that the unmeasured detection gap is often larger than the measured repair time. The article on manufacturing KPIs covers the broader KPI family this metric belongs to.

What MTTD measures

Mean time to detect is the average duration between the actual onset of a fault and the moment the maintenance team becomes aware of it. The onset can be defined in three ways depending on the asset:

  • Sensor-based, the moment a tracked parameter (vibration, current, temperature) crossed the threshold that retrospectively signalled the failure.
  • Quality-based, the moment reject rate or dimensional variance moved out of normal range.
  • Production-based, the moment cycle time, throughput rate, or micro-stop frequency drifted enough to indicate degradation.

The "awareness" end of the measurement is the timestamp of the work order opened, the alert acknowledged, or the operator's verbal flag, whichever came first.

For many assets MTTD is measured retrospectively because no leading indicator was set up in advance. The first time the metric is calculated, the team is usually looking back at the most recent failures and computing the gap from data that was always present but unwatched. The piece on root cause analysis covers the technique of reading historical data for the signal that should have triggered an earlier alert.

Why cutting MTTD often beats cutting MTTR

MTTR improvements usually require physical or procedural changes: pre-staging spares, tool drops near the asset, redesigned access panels, faster diagnostic procedures. Each of these is real work and costs real time.

MTTD improvements are usually configuration changes: setting an alert threshold on an existing sensor, routing an alert to the right person's phone instead of an unread email inbox, clustering small signals into a meaningful alarm. Same gain in avoided downtime; much smaller investment.

The math is clean on assets where degraded operation costs production minutes. If a 60-minute MTTD gap means 60 minutes of asset running at 80% throughput before the stop, that is 12 minutes of production-equivalent loss that is currently invisible.

Cutting MTTD to 15 minutes recovers 9 of those 12 minutes per event, and the asset stops sooner, usually with a smaller repair scope. The article on the preventive maintenance schedule covers how MTTD reductions feed back into PM cadence decisions.

Where most plants find MTTD reductions

1. The signal exists but nobody is watching it

The most common case. A sensor is installed and recording, but no alert is configured. The data is there in the historian; the alert layer is missing. Configuring an alert on an existing data stream is usually a half-day of work and a same-day gain.

2. The signal is being aggregated past usefulness

Some plants average raw data over an hour or a shift, which smooths out the short-duration excursions that are the early warning of failure. The fix is to keep a high-resolution stream for alert purposes even if the aggregated stream is used for reporting.

3. The alert exists but is routed wrong

An alert that pings an unread mailbox or fires to a control-room screen no one watches is not an alert. Routing has to match the operating reality, to the technician on duty's phone, to the line supervisor's tablet, to a clear physical signal at the asset.

4. Cluster signals are ignored

Three micro-stops on the same asset in one shift is a stronger signal than any single stop. Most alert systems do not pattern-match across multiple events; configuring a cluster rule moves detection from "asset has stopped" to "asset is about to stop." The work order management system covers cluster-rule structures in more detail.

What good MTTD numbers look like

For a well-instrumented asset class with good alert routing, MTTD typically runs 3-8 minutes, enough time for someone to notice, decide, and respond. For asset classes with no leading-indicator instrumentation, MTTD often runs 45 minutes to several hours, because detection happens only when production output drops far enough to be visible.

The targets to aim for:

  • Production-stopping critical assets: MTTD under 5 minutes.
  • Production-degrading critical assets: MTTD under 20 minutes.
  • Non-critical assets: MTTD acceptable up to 4 hours, because the cost of constant monitoring exceeds the failure cost.

The thresholds match asset criticality; tighter monitoring on lower-criticality assets is over-investment.

How Fabrico fits

The MTTD concept works in any plant with sensors and an alerting layer.

Where a unified OEE + CMMS platform helps is in two places: the leading-indicator data and the work-order data live in the same system, so MTTD can be calculated retrospectively from event timestamps without a manual reconciliation, and the alert routing rules can be tied to asset criticality from the CMMS hierarchy.

Fabrico is built so MTTD is a tracked metric alongside MTBF and MTTR rather than an afterthought. To see what your MTTD picture looks like across your critical asset classes, book a demo .

Frequently asked questions

How do we calculate MTTD on assets without sensors?

Use production-data proxies: micro-stop frequency, reject rate, cycle-time drift. All three can be derived from the OEE event stream. The MTTD calculation against these proxies is approximate but useful.

What if the alert noise becomes too high?

That is the most common reason MTTD programs decay. The fix is threshold tuning: every false alert is a data point that the threshold needs adjustment, and the alerting team should review thresholds monthly for the first six months. A noisy alert system loses operator attention within weeks.

Should we report MTTD plant-wide?

By asset class, not plant-wide. Aggregating MTTD across critical and non-critical assets produces a meaningless number because the targets differ by an order of magnitude. Per asset class, the number is decision-relevant.

Who owns the MTTD metric?

The reliability engineer where the role exists; the maintenance manager otherwise. The owner is responsible for the alert configuration, the routing, and the monthly threshold review.

What is the most common implementation mistake?

Buying more sensors before configuring alerts on the existing ones. Most plants have more sensor data than they have alert rules. The first MTTD gain is usually in the configuration layer, not in the instrumentation layer.

Latest from our blog

Define Your Reliability Roadmap
Validate Your Potential ROI: Book a Live Demo
Define Your Reliability Roadmap
By clicking the Accept button, you are giving your consent to the use of cookies when accessing this website and utilizing our services. To learn more about how cookies are used and managed, please refer to our Privacy Policy and Cookies Declaration