What To Do in the First Hour After an Unexpected Equipment Failure
Unplanned downtime costs an average of $25,000 per hour in industrial facilities, and the decisions made in the first 60 minutes either contain that cost or compound it. The first hour isn’t about fixing the problem. It’s about securing the situation, preserving diagnostic evidence, and making the fastest possible transition from reactive scramble to directed response.
Step 1: Secure the situation
Isolate the failed equipment using lockout/tagout procedures before anyone starts troubleshooting. Assess for secondary hazards, such as fluid leaks, electrical exposure, and runaway downstream processes. If the failure is on a line with dependent equipment, determine whether adjacent machines need to be stopped. A technician working on an unsecured machine turns one failure into two incidents.
Step 2: Get eyes on it
The first look at a failed machine is the only opportunity to observe it in its failure state before anyone resets, reconnects, or disturbs the configuration. Photograph the fault display, the physical condition of components, and anything visibly wrong. Note what was happening when the failure occurred, including the speed, load, temperature, and cycle stage. This evidence disappears the moment the machine gets touched.

Step 3: Pull the fault log
Most drives and PLCs store motor current, bus voltage, output frequency, and run time at the time of the trip. Resetting the unit erases this on many platforms. Pull the fault history before touching the reset button. A fault code with context is a starting point. A cleared fault log is a blank page.
Step 4: Check inventory and lead times
While the diagnosis is underway, check what’s on the shelf. Identify whether a spare exists in inventory, whether repair is faster than replacement, and what the lead time is on a new unit if neither is available. This information shapes the recovery timeline before the timeline shapes itself.
Step 5: Make a repair decision
The factors that should inform a repair decision include the age of the failed component, whether the failure is component-level or catastrophic, spare parts availability, repair turnaround time, and whether this failure has occurred on the equipment before. A component-level failure on a 4-year-old drive is a different decision than a catastrophic failure on a 12-year-old one.

Step 6: Document thoroughly
A documented fault code and operating conditions at failure will get a technician started faster than “it stopped working.” GES’s emergency repair line is available 24 hours a day, 365 days a year. Having that number on hand before an emergency is the preparation that makes this step fast.
Step 7: Log everything
Log the fault code, physical symptoms, operating conditions, the failed component, and the repair action taken. A CMMS entry that captures all of this turns a reactive repair into institutional knowledge. The next technician who works on this equipment starts with context rather than zero.
The scramble is inevitable, but the structure isn’t
The decisions made in the first hour — what to preserve, what to check, who to call, and with what information — either shorten the downtime or extend it. Most facilities get faster at this through repetition. The ones that build the structure before the failure don’t have to contend with the pressure of a ticking clock the moment a machine goes down.