Every plant has a list of machines everyone complains about. The capper on line 2. The old filler. The palletizer that nobody wants on their shift.
The list is usually wrong.
It was built from breakdowns: the machines that stop dramatically, get a maintenance call and a reason code, and come up in the morning meeting. Those are the loud machines. The costly ones, the machines that quietly take a few percent of output every shift without ever breaking down, are not on it, because nobody experiences them as a problem. And the ranking is in downtime minutes, which is not the same thing as lost output, which is not the same thing as lost money.
This article is about replacing the folklore list with a measured one. It takes station-level stop data, joins it to the maintenance history, and produces two things: a ranked list of the assets actually responsible for most of the lost output, and a four-quadrant matrix that tells the plant manager and the maintenance manager what to do about each of them.
A bad actor is an asset whose losses are out of proportion to its share of the line. Not the oldest machine, not the one with the most work orders, not the one maintenance dislikes: the one that removes the most output per week relative to everything around it.
Pareto applies brutally here. On multi-line plants we have measured, 15 to 20% of machines account for 70 to 80% of lost output. The rest are noise. That concentration is the opportunity: a plant does not have to improve sixty machines, it has to improve ten, and it has to know which ten.
The definition has to be in lost output, not downtime minutes. The next section is about why that distinction decides whether the list is useful.
Downtime minutes. This is what the maintenance system gives you: hours of recorded stops per asset. It is the basis of every folklore list, and it is wrong for a simple reason. An hour down on a machine with an accumulation table behind it may cost the line nothing; the buffer absorbs it. An hour down on the bottleneck costs the whole line an hour. Downtime minutes treat these as equal.
Lost output. Downtime weighted by whether line output actually fell while the machine was stopped. This is the right basis for finding bad actors, and it requires two things the maintenance system does not have: stop data at station level, and the line’s count at the same moments. A machine with four hours of stops and no effect on output is not a bad actor. A machine with ninety minutes of stops that each starve the filler is.
Lost value. Lost output weighted by the SKU running at the time. A stop during a high-margin format costs more than the same stop during a low-margin one. This is the right basis for deciding where to spend, and it is the one that turns a maintenance list into a finance conversation.
Most plants rank on the first basis because it is the only one their systems can produce. The gap between the first and the second is where the folklore list goes wrong.
Four biases, and every plant has all of them.
Loud versus costly. A breakdown that stops the line for two hours once a month is remembered, coded and discussed. A case packer that stops for twenty seconds every three minutes is cleared by the operator and forgotten. The second one loses more output. The micro-stops article is about why the second kind is invisible; here the point is that it never makes the list.
Recency. The machine that failed last week is at the top of everyone’s mind and therefore at the top of the list. The machine that has lost 3% of output every week for a year is nowhere, because nothing about it was ever an event.
The machine maintenance hates. Some assets are miserable to work on: awkward access, poor documentation, a vendor who does not answer. They generate complaints and long work orders, and they are remembered as bad actors whether or not they cost the line anything.
The non-bottleneck trap. The “worst” machine on many lines sits upstream of a buffer. Its stops are absorbed, and the real cost appears on the bottleneck downstream, which shows as “starved” and gets blamed for nothing. Unless stops are attributed to the machine that caused them rather than the one that showed them, the bottleneck looks fine and the upstream machine looks worse than it is, or the other way round.
None of these biases is foolish. They are what a ranking looks like when it is built from memory and work orders instead of from output data.
Five steps. The first three are OEE data; the fourth is maintenance data; the fifth is the join.
One precondition, and it is the one that stops most plants. Step five only works if the OEE system and the maintenance system refer to the same machine by the same name. In plants where the stations were set up by the operations team and the assets by the maintenance team, years apart, they usually do not. On one group’s plants, the first attempt at the join matched roughly three quarters of assets at one site and could not be done at all at another, because the naming conventions had nothing in common. The matching exercise is the first deliverable, not a detail, and it is worth saying so in the project plan. A plant that fixes the naming once has solved a problem that will otherwise block every piece of cross-system analysis it ever tries.
Plot every machine on two axes: lost output from the OEE data, maintenance cost from the CMMS. Four quadrants appear, and each one has a different right answer. This is the chart the plant manager and the maintenance manager can finally read together, because it uses both of their numbers.
| Low maintenance cost | High maintenance cost | |
|---|---|---|
| High lost output | Under-maintained, or not a maintenance problem. The machine loses output and nobody is spending on it. Either the PM is missing the failure mode, or the cause is operations: settings, material, operator habits, format parts. Check the stop pattern before adding maintenance. | The true bad actor. Loses output and consumes budget. Root-cause analysis on the top three stop reasons, PM redesign against the failure modes the data shows, and if that fails, the replacement case, built on lost value. |
| Low lost output | Leave alone. It works and it is cheap. Resist the urge to improve it. | Over-maintained, or a parts problem. The spend may not be buying output. First check that the PM is not what keeps the machine running well, by reviewing it against the failure modes it is meant to prevent. Then extend intervals, review the spares strategy and question the service contract. This quadrant funds the work in the other two. |
The lower-right quadrant is the one plants forget. Maintenance budgets are rarely cut from assets that are running well, so money sits on machines that would run just as well with half the attention, while the upper-left quadrant starves. Moving spend from lower-right to upper-left, once the PM review confirms it is safe, is usually the cheapest capacity gain on the list, and it is invisible without the matrix.
Take the five machines at the top of the lost-output ranking. For each one, the data should already have narrowed the cause; the work is to act on it.
Root-cause the top three stop reasons. Not all stop reasons on the machine, the top three. They usually carry most of its loss, and they are usually fewer than the plant expects.
Check the PM against the failure modes. Most preventive maintenance schedules were written from the OEM manual and have never been compared with what actually fails. If the data says the machine stops for sensor faults and the PM is about lubrication, the PM is well-executed and useless. Rewrite the tasks against the stop reasons.
Add condition monitoring where the pattern is wear. A stop rate that climbs steadily between interventions is a wearing component. Monitor the condition and intervene before the stop, rather than on a calendar.
Build the upgrade case on lost value. For the machine that is structurally too slow or too unreliable, this is the moment the capex argument is finally made with data: lost value per year, against the cost of the upgrade. It is a much smaller and better-targeted request than a new line.
Or accept it and buffer it. Some bad actors are cheaper to live with. If a small accumulation table removes the effect on the bottleneck, that may be the whole fix.
Re-measure at 90 days. The Pareto should have changed shape: the top five should have dropped, and a new top five, with smaller losses, should have appeared. If the same five are still there, the actions did not address the cause, and the data will say which ones.
Take the six-line plant used across this series, about sixty machines in total. Ten weeks of station-level data, attributed and weighted, produced a Pareto in which eleven machines carried 76% of lost output. The other forty-nine shared the remaining quarter.
The top five, with where they landed on the matrix:
| Rank | Machine | Lost output share | Maintenance cost | Quadrant | What the data showed |
|---|---|---|---|---|---|
| 1 | Case packer, line 3 | 14% | Low | Under-maintained / operations | Chronic micro-stops on two case formats, climbing after changeover. Settings and a worn change part. Not on the folklore list. |
| 2 | Filler, line 1 | 12% | High | True bad actor | Sensor and valve faults; PM schedule was about lubrication. PM rewritten against the stop reasons. |
| 3 | Palletizer, shared | 11% | Medium | True bad actor in effect | Looked harmless in isolation; blocks three lines when it stops. Every stop attributed back to it was a line stop. |
| 4 | Labeler, lines 2 and 3 9 | % L | ow U | nder-maintained / operations S | ame format problem on two lines. One fix, two machines. |
| 5 | Depalletizer, line 5 | 7% | High | True bad actor | Long repair times; spares not stocked. MTTR was the problem, not MTBF. |
The capper on line 2, the machine everyone would have named first, ranked ninth. It broke down dramatically about once a month and had a large accumulation table behind it, so most of its downtime never reached the line count. It stayed on the maintenance plan and came off the priority list.
At 90 days the top five had dropped to a combined 31% of lost output from 53%, and the Pareto had a new, flatter top. The numbers are illustrative, but the shape is what we see: a chronic machine nobody blamed at the top, a shared downstream asset that looked innocent, and a famous bad actor that was mostly loud.
Maintenance budgets are usually allocated by asset age, OEM schedule and last year’s spend. None of those is lost output. The result is a plan that spends evenly across machines that lose very different amounts of capacity.
The bad-actor list reallocates by lost value: more on the five, less on the forty, with the lower-right quadrant of the matrix funding the move. It also gives the maintenance manager the argument they have rarely been able to make to finance: this spend on this machine returns this much output, measured, with a 90-day re-check. Maintenance stops being a cost line that gets trimmed each year and becomes a capacity investment with a return, which is what it always was.
A traditional CMMS is a system of record for maintenance events. It knows every work order, every part issued and every hour of labor. It can rank machines by work-order count, by downtime minutes and by maintenance cost. What it cannot know is what any of that cost the line in output, because it has never seen the line’s count. The downtime-minutes list is the best a CMMS can do, and the first half of this article is about why that list is wrong.
Fabrico’s manufacturing performance platform (MES, OEE, CMMS & AI) was built with OEE, maintenance, production scheduling and SKU-level output value in one system, on one asset model, and the bad-actor analysis is where that pays off. Each module adds a piece a CMMS does not have.
One asset, two histories. The station the OEE module measures is the asset the maintenance module maintains. There is no join to build, because there are not two systems to join. The naming problem that blocks most plants at step five does not arise, and where a plant arrives with an existing CMMS the matching is done once, during onboarding, rather than by every analyst who tries later.
Lost output, not downtime minutes. Every stop is captured at station level by sensor, PLC signal or computer vision, attributed to the machine that caused it through the starvation-and-blocking logic, and weighted by its effect on line output. The ranking is on the right basis by default. A machine with hours of absorbed downtime does not appear at the top; the palletizer that quietly blocks three lines does.
Lost value, using the plant’s own SKU economics. With the selling price and margin entered per SKU, each machine’s lost output is valued at the margin of what was running when it stopped. The Pareto and the matrix are drawn in dollars as well as hours, which is the version finance will read.
The schedule knows which machines to trust. This is the piece that turns the list into shop-floor behavior. The production scheduling module plans against each line’s demonstrated performance rather than its nameplate, so a known bad actor is scheduled at the output it actually delivers, not the output the standard assumes. It routes the high-margin SKUs to the reliable lines and the forgiving ones to the line with the chronic case packer. It puts the PM window for a bad actor into the plan as a block, instead of the planner discovering it as a surprise. And when the bad actor stops anyway, the replanning engine re-sequences the affected orders across the other lines within minutes, with the financial impact module showing what the stop just cost and what the replan recovers. The maintenance decision, the scheduling decision and the financial consequence are the same event seen from three sides.
The matrix as a standing view, not a project. Lost output from the OEE side and cost from the maintenance side are the two axes of a chart that updates as the data does. The quadrant a machine sits in this month is visible without anyone exporting a spreadsheet.
AI insights that rank the fix, not just the machine. The platform’s actionable insights do not stop at “the case packer on line 3 is a bad actor”. They propose the intervention the stop pattern points to, whether that is a PM task rewritten against the real failure modes, a format-part replacement, a condition-monitoring trigger or an operations fix, and rank the proposals across the whole plant by recoverable value. The plant decides; the system proposes and then measures whether the Pareto moved.
The fix becomes a work order in the same system. When the action is maintenance, it is raised as a work order against the same asset, with the lost-output evidence attached, and PM compliance on that asset is tracked back against the stop rate. The loop closes without a handoff between tools.
The difference from a CMMS is not a better maintenance system. It is that maintenance, scheduling and value are modules of one platform whose unit of account is production output, so a bad actor is at once a maintenance priority, a scheduling constraint and a line on the P&L, and the plant can act on all three without reconciling three systems.
What is bad-actor analysis in manufacturing? Ranking equipment by the production output it causes the line to lose, rather than by downtime minutes or work-order count, and then sorting those assets by maintenance cost to decide which need root-cause work, which need a different kind of attention and which should be left alone.
How many machines are usually bad actors? On multi-line plants, 15 to 20% of machines typically account for 70 to 80% of lost output. On a sixty-machine plant that is ten to twelve assets.
Should I rank machines by downtime or by lost output? Lost output. Downtime minutes on a machine with a buffer behind it may cost the line nothing, while a shorter stop on the bottleneck costs the whole line. Downtime minutes are what a CMMS can produce; lost output needs station-level stop data joined to the line count.
How often should the bad-actor list be refreshed? Re-measure at 90 days after acting on the top five, and then quarterly. If the same machines are still at the top, the actions missed the cause.
Is a bad actor a maintenance problem or an operations problem? The matrix answers that per machine. High lost output with high maintenance cost is a maintenance problem. High lost output with low maintenance cost is usually operations: settings, material, format parts or operator habits. The mistake is assuming all bad actors belong to maintenance.
If the plant’s list of problem machines was built from breakdowns and memory, it is worth testing against output data. The test is cheap and the result is usually a different list.
The fixed-scope pilot does it on one line: six weeks, station-level stop capture with an industrial sensor and hub installed with your team, stops attributed to the machine that caused them and weighted by effect on output, and a readout with the line’s own bad-actor ranking and the three highest-value interventions. Where the plant has a maintenance history for the same assets, the matrix is in the readout too. The fee is fixed and credited in full against a first-year subscription if you roll out. The pilot runs on Fabrico’s manufacturing performance platform (MES, OEE, CMMS & AI), which connects machine data, OEE and loss analysis, production scheduling, SKU-level output value and maintenance in one system, so the bad-actor list stays current as the fixes land.
Request a demo or read how the maintenance module shares one asset model with OEE.