Maintenance has its KPIs: PM compliance, work-order backlog, MTTR, wrench time. Reliability has its own: MTBF, failure rate, availability by asset. Operations has OEE.
Three sets of numbers, reviewed in three meetings, and in most plants nobody has ever shown that the first two move the third. PM compliance is 95% and availability is falling. MTBF on the filler is improving and the line is still short. The maintenance manager is hitting every target and the plant manager is still buying a new line.
This article is about two maintenance numbers that do predict availability, roughly a month ahead: PM compliance on the assets that matter, and the rate at which fixes actually hold. It is also about why proving that requires OEE data and maintenance data on the same machine, which is the thing most plants’ systems cannot do.
The maintenance system counts tasks. The OEE system counts output. Neither knows what the other saw.
A CMMS can report 95% PM compliance. It cannot tell you that the 5% missed were the ten machines responsible for most of the lost output, or that the PMs completed on time address lubrication while the machines fail on sensors. Backlog can be clear while the same valve is replaced four times in a quarter, because each replacement closes its own work order. MTBF can improve on every asset while line availability falls, because MTBF counts failures per asset and a palletizer that fails rarely but blocks three lines when it does is worth more than ten fillers that fail often with a buffer behind them.
The reliability metrics are not wrong. They are asset-level and output-blind. Without a weighting by what each asset’s downtime cost the line, MTBF ranks machines the same way downtime minutes do, which the bad-actor article covers: loud machines first, costly ones nowhere.
The link between maintenance and output cannot be seen without joining the two datasets on the same machine. That join has a precondition, stated once here and in the bad-actor article: the OEE stations and the CMMS assets have to share a naming convention. Where they do not, matching them is the first job. On one group’s plants, the first attempt matched about three quarters of assets at one site and none at another.
Once the data is joined, two maintenance numbers turn out to lead availability by two to six weeks.
PM compliance on the bad actors, weighted by lost output. Not plant-wide compliance. Take the ten or twelve assets that carry most of the lost output, and measure compliance on those. Then weight it: a missed PM on the asset that loses 14% of output counts fourteen times a missed PM on one that loses 1%. A plant at 95% plant-wide and 60% on the top ten, weighted, is a plant at roughly 60%, and its availability will show it a month later.
Fix-held rate. The share of corrective work orders after which the same failure mode does not recur on the same asset within 30 days. It is the honest measure of whether maintenance fixed something or reset it.
Fix-held rate = (corrective work orders with no repeat of the same reason code on the same asset within 30 days) ÷ (all corrective work orders)
A fix-held rate of 40% on a valve means the valve was “fixed” and failed again six times in ten. MTBF on that asset counts six failures and improves slightly each time the interval stretches; fix-held says the problem was never solved. It is the complement of the repeat-failure rate that reliability engineers already track: a 40% fix-held rate is a 60% repeat-failure rate. Many CMMSs can report repeat work orders, but few can compute it from the line, because they do not link a repeat stop to the prior work order on the same asset; the repeat is just a new work order.
MTBF stays in the picture, but weighted by lost output and read alongside fix-held. On its own it tells you how often an asset fails. Weighted, it tells you how much that costs. Paired with fix-held, it tells you whether maintenance is reducing the failures or recycling them.
The typical shape we see once the join exists: on the same line, assets with fix-held above 80% and weighted PM compliance above 90% run 8 to 12 availability points higher than assets below both. The two groups are usually maintained by the same team to the same plant-wide compliance target. The difference is in which PMs were missed and which fixes held.
Both metrics lead. Availability lags. That is what makes them useful.
A missed PM on a bad actor shows up first as a rising stop rate on that asset, two to six weeks later, as the component the PM would have addressed begins to wear: more micro-stops, longer clears, a sensor that needs resetting more often. The micro-stops article describes this trend as a condition-based maintenance trigger. Then comes the breakdown that gets a reason code and a work order. Then availability drops in the monthly report, by which time the cause is two months old.
A low fix-held rate shows up as the same reason code returning on a cycle: valve fault, work order, valve fault, work order. On a per-asset timeline it is unmistakable, a sawtooth of repairs with the stop rate climbing between them. In the CMMS it is six closed work orders and a reasonable MTTR.
Draw the timeline for one bad actor: PMs scheduled, PMs done or missed, corrective work orders, and the station’s stop rate per week. The stop rate climbs after each missed PM and resets after each fix that held, and keeps climbing after each fix that did not. That one chart is the evidence that maintenance moves output, and it cannot be drawn from either system alone.
The diagnosis comes from OEE; the cure is administered through the CMMS. Between them is a chain of hand-offs, and the loop breaks at whichever one is not made.
In most plants these seven steps cross two systems and at least three people, and each crossing is a place where the loop quietly stops: the failure modes never reach the PM planner, the repeat is never recognized as a repeat, the availability number is never read against the maintenance one.
Four changes to the monthly pack, none of which need new people.
The filler on line 1, bad actor number two in the bad-actor article: 12% of the plant’s lost output, high maintenance cost, firmly in the true-bad-actor quadrant.
| Before | After | |
|---|---|---|
| PM compliance on the filler | 97% | 98% |
| PM content | Lubrication, belt tension, guarding check | Sensor cleaning and alignment, valve inspection and replacement at measured interval, plus the original tasks |
| Top stop reasons | Sensor fault, valve fault, product-side jam | Product-side jam, sensor fault (reduced), other |
| Fix-held rate on valve work orders | 40% (same valve replaced four times in a quarter) | 85% (correct-spec valve, stocked as a spare) |
| Stop rate per week | 31 | 15 |
| Filler availability | 81% | 90% |
| Contribution to line OEE | About 2 points recovered from the 8 unplanned-stop points |
What changed. The PM was 97% compliant and almost entirely irrelevant: the filler did not fail on lubrication. Rewriting it against the two measured failure modes took a morning. The valve had been replaced with a near-equivalent part from stock four times; the correct spec was identified from the failure pattern, stocked, and fitted once. Fix-held on that failure mode went from 40% to 85%, the stop rate halved inside six weeks, and the filler’s availability rose 9 points. Line 1 shares the series line’s loss profile, so that was roughly 2 OEE points: the targeted-fixes allocation for unplanned stops in the hidden capacity method.
Nothing about this was expensive. The expensive part had been the two years of 97% compliance on the wrong tasks, and the only reason it was found is that the filler’s stops and its work orders were finally on the same screen. The figures are illustrative; the shape, high compliance on irrelevant PMs and a low fix-held rate on the real failure mode, is the most common finding when the join is made.
A standalone CMMS can report 97% compliance on the filler and never know the filler is losing output. That is not a flaw in the CMMS; it is what a system of record for maintenance tasks is. The loop above needs maintenance and output in one place, which is how Fabrico’s manufacturing performance platform (MES, OEE, CMMS & AI) is built.
One asset model. The station the OEE module measures is the asset the maintenance module maintains. Every stop has a work-order history and every work order has a stop history, without a join project.
Fix-held is computed, not estimated. A repeat stop on the same asset with the same reason code within the window is linked automatically to the prior corrective work order. Fix-held rate exists per asset, per failure mode and per technician from the first month.
PM compliance is weighted by lost output. Compliance is reported on the bad actors, weighted by what each asset costs the line, next to the plant-wide figure. The number that is managed is the number that matters.
Stop rate since last PM is a trigger. The climbing stop-rate trend on a bad actor raises the PM early, as a condition-based task rather than a calendar one, before the breakdown that would have earned a reason code.
The AI insights propose the PM rewrite. When an asset’s measured failure modes do not match its PM tasks, the actionable insights propose the revised task list, rank it against the rest of the plant’s backlog by recoverable value, and measure whether the stop rate moved after it was adopted. The plant decides; the system proposes and then checks.
The schedule and the financial impact module close the loop from the other side. The production scheduling module places the PM window for a bad actor in the plan as a block, so it is not skipped because the line was busy. With the selling price and margin entered once per SKU, the financial impact module shows what a missed PM cost on the SKUs that ran afterwards, which is the number that ends the argument about whether the window could be found.
The difference is not a better maintenance system. It is that maintenance is one module of a platform whose unit of account is production output, so “did the fix hold” and “did the PM get done” are answered in the same currency as “did the line make the plan”.
What is fix-held rate? The share of corrective work orders after which the same failure mode does not recur on the same asset within a set window, usually 30 days. It measures whether maintenance fixed the problem or reset it, and it requires repeat stops to be linked to the prior work order, which few CMMSs do. It is the complement of the repeat-failure rate.
Does PM compliance improve OEE? Only when the PMs address the failure modes the asset actually has, and only on the assets that lose output. Plant-wide compliance on OEM-manual tasks can be high while availability falls. Compliance weighted by lost output, on the bad actors, predicts availability about a month ahead.
Why is PM compliance high but downtime still rising? Usually one of three reasons: the missed PMs are concentrated on the bad actors, the PM tasks address the wrong failure modes, or corrective fixes are not holding and the same failure recurs. The joined OEE-and-CMMS data shows which.
What maintenance KPIs predict availability? Weighted PM compliance on the top lost-output assets, fix-held rate, and stop rate since last PM. MTBF and MTTR describe what happened; the three above lead it by two to six weeks.
How do you link CMMS data to OEE? The OEE stations and the CMMS assets have to refer to the same machine by the same identifier. Where they do not, match them once, then attribute every stop to its asset and every work order to the stop it addresses. A platform with one asset model for both removes the join entirely.
If maintenance is hitting its targets and availability is still falling, the targets are measuring the wrong things on the wrong assets. Joining the stop data to the maintenance history for the line’s ten bad actors will show which, usually in the first month.
The fixed-scope pilot does exactly that on one line: six weeks, station-level stops captured with an industrial sensor and hub installed with your team, attributed to the causing asset, and matched to the plant’s maintenance history where asset names allow. The readout includes the line’s bad actors, their measured failure modes against their current PM tasks, fix-held rate on the recurring failures, and the three highest-value interventions, which on most lines include one PM rewritten against what actually fails. The fee is fixed and credited in full against a first-year subscription if you roll out. The pilot runs on Fabrico’s manufacturing performance platform (MES, OEE, CMMS & AI), which connects machine data, OEE and loss analysis, production scheduling, SKU-level output value and maintenance in one system, so maintenance and output are read from the same asset history.
Request a demo or read how the maintenance module shares one asset model with OEE.