Key takeaways
See our roundup of RCM software for prioritizing what matters first.
Most reliability engineers arrive at a new plant with energy, a methodology, and an opinion about what the team should be doing. The maintenance team has seen this arrival before, often more than once. They are polite, slightly skeptical, and watching to see whether this reliability engineer will do the same thing as the last one: propose a program, run a few interventions, leave a partial implementation behind.
The pattern that breaks this cycle is not new methodology. It is a deliberate diagnostic period, the first 30 days spent not fixing anything visible, but building a defensible view of what the failure picture actually looks like. That view becomes the reliability engineer's credibility currency for the next two years.
The first week is not for data analysis. It is for understanding how the maintenance team thinks about the plant. Three conversations every reliability engineer should have in week 1:
These conversations are diagnostic. The reliability engineer is not solving anything yet; they are learning whose mental model of the plant they are about to challenge with data.
The CMMS holds the failure history. Most of it is messy. Free-text reason codes, inconsistent asset names, work orders that close without a documented failure mode. The first cleanup pass is just standardising the failure-mode field across the last 12 months of work orders. The article on the work order management system covers the data structures this depends on.
Do not try to clean everything. The first cut covers the top 20 assets by work-order count. That subset typically accounts for the large majority of failure events and is enough to anchor the next 60 days.
By end of week 4, the reliability engineer has one document: a table of the top 20 assets with a normalised failure-mode count and an estimated production-impact per failure. The document is not a recommendation. It is a baseline.
The reliability engineer presents this baseline at the end of week 4 to the maintenance manager and the production manager, not to the plant manager, not to a steering committee. Two people. The point of this meeting is to validate that the data reflects what the room sees on the floor.
If it does not, the data is wrong and another two weeks of cleanup is needed. If it does, the next phase can start. The piece on production loss analysis covers the production-impact estimation.
The reliability engineer now has a credible failure-history view. The next step is to cross it against the OEE loss data. An asset with many failures but low production impact is a different problem from an asset with few failures but high production impact. Without the cross, the prioritisation goes to the loud cases rather than the costly ones.
The output is a ranked list of three asset classes that together account for most of the avoidable unplanned production loss. Three is the right number, fewer means the analysis missed something, more means the prioritisation is too diffuse to act on. The piece on manufacturing KPIs covers the underlying ranking logic.
Of the three, pick the asset class with the best ratio of impact-to-fixability. Spend two weeks on this one class: pull every work order from the last 12 months, walk the assets, interview the technicians who repaired them, identify the dominant failure mode.
The output is a one-page failure-mode analysis with three to five proposed interventions (PM change, parts upgrade, training, design change) and an estimated impact on production minutes per quarter.
This is not yet an intervention. It is an option set, with numbers attached.
From the three to five options, pick the one with the highest impact-to-effort ratio. Resist the urge to combine multiple interventions; the first one needs to be measurable in isolation. Most plants over-pick the scope of the first reliability intervention and end up unable to attribute the gain.
The intervention typically falls into one of three buckets: a tightened PM frequency on a specific failure mode, a parts upgrade on one asset class, or a procedural change to how the work order is executed. All three are achievable in 4-6 weeks. The framing connects to root cause analysis at the failure-mode level, the intervention only works if it is matched to the cause, not the symptom.
The intervention runs. The reliability engineer is on the floor for the first execution, not in the office. The measurement plan is set in advance, leading indicator (compliance with new PM, parts upgrade complete) and lagging indicator (failure count on that mode, production minutes lost) with explicit targets.
By day 90, the lagging indicator has not moved yet, most reliability interventions take a few months to show up in lagging metrics. The leading indicator should be at or near target. The reliability engineer presents both numbers to the plant manager, with an explicit "the lagging gain will arrive over the following quarter" note.
The piece on the preventive maintenance schedule covers how the new PM cadence becomes a standing rule rather than a one-time change.
Not a transformation. Not a glossy program document. The output is:
The next 90 days are easier because of those four outputs. The reliability engineer who skips this opening phase to start with action usually spends the second quarter recovering credibility from a misfired intervention.
The first 90 days work in any plant with a CMMS and an OEE system. They work faster when the failure history and the OEE losses live in the same database under a single asset hierarchy, the cross-analysis at weeks 5-6 takes hours instead of weeks.
Fabrico is built so the reliability engineer can run that cross-analysis against the same data the maintenance team uses for daily work. To see how the first 90 days would look against your live data, book a demo .
No. A program plan presented before the data is in commits the engineer to a course of action they will need to reverse when the data arrives. The maintenance team also recognises a pre-baked plan and discounts it. Hold the plan until end of week 4 at the earliest.
Negotiate a 30-day window for the diagnostic, with a clear deliverable (the baseline document). Most maintenance managers will accept this if the deliverable is concrete and the timeline is short. The engineer who delivers the baseline on day 30 has earned the right to set the pace of the next 60 days.
The first cleanup pass is the work. Without it, every later analysis is built on sand. The article on work order management systems covers what good data looks like; the reliability engineer is the one who pulls the current data up to that standard.
Useful outcome. It means the failure-mode hypothesis was wrong, the intervention was too small, or the lagging indicator was the wrong choice. Each of those is a recoverable lesson, and the leading-indicator-first methodology means the engineer catches it at day 60 rather than day 270.
Presenting at the steering committee in week 2. The audience at week 2 should be the two people who will execute or be affected: the maintenance manager and the production manager. The plant-wide audience comes after day 90, when there is a real outcome to discuss.