Menu
The Reliability Engineer's First 90 Days at a New Plant

The Reliability Engineer's First 90 Days at a New Plant

The first 90 days for a new reliability engineer should be diagnostic, not corrective. Days 1-30 build the failure-history view, days 31-60 identify the 3.
The Reliability Engineer's First 90 Days at a New Plant

Key takeaways

See our roundup of RCM software for prioritizing what matters first.

  • The temptation in a new reliability role is to start by fixing the worst-performing asset. That almost always burns the first quarter on someone else's priority list rather than building a defensible picture of what is actually broken.
  • The right opening 90 days are diagnostic, not corrective. Days 1-30 are about getting one reliable failure-history view. Days 31-60 are about identifying the three asset classes that drive the most unplanned production loss. Days 61-90 are about shipping one defensible intervention with measurable outcome.
  • The most common opening mistake is presenting a reliability program before the data is in. The plant has heard the same program three times before; what builds credibility is the first specific, defensible number, "this asset class costs 47 hours of production a quarter", backed by the data.
  • The first 90 days do not produce a reliability transformation. They produce the trust and the data foundation that a transformation can be built on without the team rejecting it as someone else's plan.

The trap of starting with action

Most reliability engineers arrive at a new plant with energy, a methodology, and an opinion about what the team should be doing. The maintenance team has seen this arrival before, often more than once. They are polite, slightly skeptical, and watching to see whether this reliability engineer will do the same thing as the last one: propose a program, run a few interventions, leave a partial implementation behind.

The pattern that breaks this cycle is not new methodology. It is a deliberate diagnostic period, the first 30 days spent not fixing anything visible, but building a defensible view of what the failure picture actually looks like. That view becomes the reliability engineer's credibility currency for the next two years.

Days 1-30: build one reliable failure-history view

Week 1: meet the team and the assets, in that order

The first week is not for data analysis. It is for understanding how the maintenance team thinks about the plant. Three conversations every reliability engineer should have in week 1:

  • The maintenance manager, what keeps them up at night. Listen for the asset names that recur.
  • The two senior technicians, what they actually do at 2am when something fails. Listen for which assets they hate fixing.
  • The production manager, which line losses they would pay to eliminate. Listen for the gap between what production thinks the cause is and what maintenance thinks.

These conversations are diagnostic. The reliability engineer is not solving anything yet; they are learning whose mental model of the plant they are about to challenge with data.

Weeks 2-3: pull the failure history

The CMMS holds the failure history. Most of it is messy. Free-text reason codes, inconsistent asset names, work orders that close without a documented failure mode. The first cleanup pass is just standardising the failure-mode field across the last 12 months of work orders. The article on the work order management system covers the data structures this depends on.

Do not try to clean everything. The first cut covers the top 20 assets by work-order count. That subset typically accounts for the large majority of failure events and is enough to anchor the next 60 days.

Week 4: produce one document

By end of week 4, the reliability engineer has one document: a table of the top 20 assets with a normalised failure-mode count and an estimated production-impact per failure. The document is not a recommendation. It is a baseline.

The reliability engineer presents this baseline at the end of week 4 to the maintenance manager and the production manager, not to the plant manager, not to a steering committee. Two people. The point of this meeting is to validate that the data reflects what the room sees on the floor.

If it does not, the data is wrong and another two weeks of cleanup is needed. If it does, the next phase can start. The piece on production loss analysis covers the production-impact estimation.

Days 31-60: identify the three asset classes that drive loss

Weeks 5-6: cross failure history with OEE losses

The reliability engineer now has a credible failure-history view. The next step is to cross it against the OEE loss data. An asset with many failures but low production impact is a different problem from an asset with few failures but high production impact. Without the cross, the prioritisation goes to the loud cases rather than the costly ones.

The output is a ranked list of three asset classes that together account for most of the avoidable unplanned production loss. Three is the right number, fewer means the analysis missed something, more means the prioritisation is too diffuse to act on. The piece on manufacturing KPIs covers the underlying ranking logic.

Weeks 7-8: deep-dive one class

Of the three, pick the asset class with the best ratio of impact-to-fixability. Spend two weeks on this one class: pull every work order from the last 12 months, walk the assets, interview the technicians who repaired them, identify the dominant failure mode.

The output is a one-page failure-mode analysis with three to five proposed interventions (PM change, parts upgrade, training, design change) and an estimated impact on production minutes per quarter.

This is not yet an intervention. It is an option set, with numbers attached.

Days 61-90: ship one defensible intervention

Weeks 9-10: pick the smallest defensible intervention

From the three to five options, pick the one with the highest impact-to-effort ratio. Resist the urge to combine multiple interventions; the first one needs to be measurable in isolation. Most plants over-pick the scope of the first reliability intervention and end up unable to attribute the gain.

The intervention typically falls into one of three buckets: a tightened PM frequency on a specific failure mode, a parts upgrade on one asset class, or a procedural change to how the work order is executed. All three are achievable in 4-6 weeks. The framing connects to root cause analysis at the failure-mode level, the intervention only works if it is matched to the cause, not the symptom.

Weeks 11-12: execute and measure

The intervention runs. The reliability engineer is on the floor for the first execution, not in the office. The measurement plan is set in advance, leading indicator (compliance with new PM, parts upgrade complete) and lagging indicator (failure count on that mode, production minutes lost) with explicit targets.

By day 90, the lagging indicator has not moved yet, most reliability interventions take a few months to show up in lagging metrics. The leading indicator should be at or near target. The reliability engineer presents both numbers to the plant manager, with an explicit "the lagging gain will arrive over the following quarter" note.

The piece on the preventive maintenance schedule covers how the new PM cadence becomes a standing rule rather than a one-time change.

What the first 90 days produce

Not a transformation. Not a glossy program document. The output is:

  • One trusted failure-history view for the top 20 assets.
  • A ranked list of three asset classes driving most of the avoidable loss.
  • One defensible intervention shipped, with leading indicators at target.
  • The maintenance manager and the production manager both believing the reliability engineer can read the plant.

The next 90 days are easier because of those four outputs. The reliability engineer who skips this opening phase to start with action usually spends the second quarter recovering credibility from a misfired intervention.

How Fabrico fits

The first 90 days work in any plant with a CMMS and an OEE system. They work faster when the failure history and the OEE losses live in the same database under a single asset hierarchy, the cross-analysis at weeks 5-6 takes hours instead of weeks.

Fabrico is built so the reliability engineer can run that cross-analysis against the same data the maintenance team uses for daily work. To see how the first 90 days would look against your live data, book a demo .

Frequently asked questions

Should the reliability engineer present a program plan in week 1?

No. A program plan presented before the data is in commits the engineer to a course of action they will need to reverse when the data arrives. The maintenance team also recognises a pre-baked plan and discounts it. Hold the plan until end of week 4 at the earliest.

What if the maintenance manager wants action immediately?

Negotiate a 30-day window for the diagnostic, with a clear deliverable (the baseline document). Most maintenance managers will accept this if the deliverable is concrete and the timeline is short. The engineer who delivers the baseline on day 30 has earned the right to set the pace of the next 60 days.

How do we handle plants with a chaotic CMMS?

The first cleanup pass is the work. Without it, every later analysis is built on sand. The article on work order management systems covers what good data looks like; the reliability engineer is the one who pulls the current data up to that standard.

What if the first intervention does not move the lagging indicator?

Useful outcome. It means the failure-mode hypothesis was wrong, the intervention was too small, or the lagging indicator was the wrong choice. Each of those is a recoverable lesson, and the leading-indicator-first methodology means the engineer catches it at day 60 rather than day 270.

What is the most common opening mistake?

Presenting at the steering committee in week 2. The audience at week 2 should be the two people who will execute or be affected: the maintenance manager and the production manager. The plant-wide audience comes after day 90, when there is a real outcome to discuss.

Latest from our blog

Define Your Reliability Roadmap
Validate Your Potential ROI: Book a Live Demo
Define Your Reliability Roadmap
By clicking the Accept button, you are giving your consent to the use of cookies when accessing this website and utilizing our services. To learn more about how cookies are used and managed, please refer to our Privacy Policy and Cookies Declaration