
Key takeaways
A failure at 2pm has a maintenance manager, a production manager, a planner, three technicians and a supervisor within walking distance. A failure at 2am has the operator on the line, an on-call technician 45 minutes away, and a supervisor who may or may not answer the phone. The fix is rarely harder than the daytime equivalent. The decision around the fix is.
The cost shows up two ways. The first is over-escalation, every soft signal becomes a 3am phone call, the on-call technician burns out, and within a quarter the team stops escalating at all. The second is under-escalation, a real critical failure waits until the morning shift, by which point the lost production has compounded.
A runbook is the bridge. It defines, in advance, which failure types get escalated, when, and how. The on-shift operator no longer has to guess.
Any failure with a safety implication (injury, fire, smoke or burning smell, hazardous material release, damaged or bypassed guarding) or a regulatory implication (food contact, pharma containment) is a one-step escalation: make the line safe, call emergency services first if anyone is hurt or there is fire, smoke, or a hazardous release, then call the supervisor and EHS. There is no triage; there is no "wait an hour." The runbook makes this the first page so the operator never has to decide.
An asset that has stopped the line. The operator runs a 60-second triage (covered below) before calling the on-call technician. The triage either confirms the failure type or buys five minutes of context for the call.
The line is producing, but quality is degrading, reject rates climbing, dimensions drifting. The runbook says: log it, capture three samples, mark "Watching" in the handoff, and only escalate if rejects exceed a defined threshold within an hour. Many quality drifts can wait for day shift, but only if the product made during the watch hour is held until it is released. Take samples only from the normal sample point, never by reaching into a running machine. Anything touching a critical control point, foreign body detection, allergen control, or seal integrity is class 1, not class 3.
A new noise, a slightly different vibration pattern, an intermittent process warning. One exception: any warning from a safety device (guard interlock, emergency stop circuit, gas, fire, or overtemperature alarm), or any smoke, burning smell, sparking, or hazardous leak, is class 1, not class 4. These get logged and added to the next handoff under "Watching." No escalation. The article on the preventive maintenance schedule covers how those soft signals feed back into PM tuning.
Every class and every step below sits under these rules. If a rule and the runbook ever disagree, the rule wins.
Before any class-2 escalation, the operator runs a fixed triage. The goal is a useful call to the on-call technician, not a panicked one, and the triage is done from outside the guarding without opening guards or touching the machine.
These five fields go to the on-call technician in writing before the phone call. The call itself becomes a 90-second conversation about action, not a 15-minute conversation about what is happening. This is the single highest leverage step in the runbook.
Each failure class has one tree. The tree is short on purpose.
Make safe with the emergency stop or normal stop → if anyone is hurt, or there is fire, smoke, or a hazardous release, raise the site alarm and call emergency services first → call supervisor → call EHS → log to CMMS with class-1 flag. No further escalation needed; supervisor takes over.
Run 60-second triage → call on-call technician → log to CMMS with triage data attached. If on-call technician unreachable after 10 minutes, call backup. If backup unreachable after another 10 minutes, call supervisor. After 30 minutes total of no contact, default action: keep the line stopped using the normal stop procedure, keep people out of guarded areas, and document for day shift. Nobody works on the equipment unless they are authorized for that task and it is locked out.
Capture three samples → log to CMMS as "in progress, quality watching" → keep running for one hour with that hour's product on hold → if reject rate above threshold, call on-call technician. If reject rate stabilizes or improves, hand off in writing to day shift with the sample data attached. The article on root cause analysis covers the sample-capture protocol that makes day-shift investigation possible.
One-line note in handoff under "Watching." No escalation. Reviewed at morning standup.
Most of the triage data is already in the system if the plant runs a unified MES and OEE platform with CMMS built in. Asset ID, time of failure, prior "Watching" entries, recent OEE events on the asset, all available without the operator typing. The runbook becomes a structured wizard at the asset station, not a paper form.
Part of the escalation tree can run through the system: a class-2 entry in the CMMS can notify the on-call technician with the triage data attached; if nobody responds within 10 minutes, the operator moves to the backup. The operator's job becomes confirming the classification and triggering the flow, not running the protocol from memory. See work order management systems for the underlying mechanics.
The risk in a night-shift runbook is over-engineering it on day one. Classes 3 and 4 are tempting to fragment into 12 subcategories; resist this. Four classes, one tree per class, one triage. Anything more complex collapses in week three when the night operator is tired.
A realistic adoption curve:
The runbook works on paper. What changes when the OEE and CMMS systems live in one platform is that the asset, the time of the stop, and recent OEE events are already in the record, the "Watching" history from earlier shifts is one tap away, and the right people can get a push notification on their phone without the operator having to look up numbers in the dark. Keep the phone call in the tree as well, because a push notification is not guaranteed to wake a sleeping technician.
Fabrico is built so the night-shift operator and the on-call technician see the same row at the same time. To see a runbook structured against your line, book a demo .
Many night shifts have someone working alone. HSE guidance says employers must manage the health and safety risks before people work alone. Build that into the runbook: a fixed check-in interval, a named person who raises the alarm if a check-in is missed, and a rule that nobody attempts class-2 work alone in an area where they cannot be seen or heard.
Write a reset rule into class 2. An operator may acknowledge and reset a fault once, and only if the procedure for that machine allows operator resets. If it trips again, the line stays stopped and the call goes out. Repeated resets hide the fault from the technician and can turn a minor trip into real damage.
Test the escalation tree monthly with a real call at night. Check that every number on the list answers, that the backup knows they are the backup, and that the list was updated after the last staffing change. An escalation tree with a dead number fails exactly when it is needed.
The classes and the triage are universal. The escalation tree (who to call) is site-specific. Plants with multiple sites should keep the structure identical and parameterize only the contact list. Inconsistent classifications across sites make group-level rollups useless.
That is the most useful disagreement to log. Every reclassification (operator said class 2, technician calls it class 3) is a training signal. Two reclassifications on the same operator usually mean the runbook description for that class needs sharper wording, not that the operator is wrong.
Two numbers: average minutes from failure to first contact with the on-call technician (target: under 5 minutes for class 2), and the rate of "wrong class on first call" (target: declining month over month). A runbook that never triggers reclassification is suspiciously perfect, usually it means class 3 and 4 are over-escalating into class 2.
No. Weekend on-call is just a different contact list. The classes and the triage are the same. Plants that maintain two separate runbooks usually let one drift while updating the other.
The "Watching" column going stale because no one reviews the handoff entries. If the soft-signal log is not read by day shift the next morning, operators stop putting effort into it within a quarter, and the early-warning channel collapses. The runbook depends on the handoff being read.
Programați o întâlnire individuală cu experții noștri sau înscrieți-vă direct în planul nostru gratuit.
Nu este nevoie de card de credit!