GuideAug 29, 2026·7 min read·

After a Major Breakdown: What to Do in the First 30 Days

The main line went down hard, everyone worked a brutal week, and things are finally running again. What you do in the next 30 days decides whether that failure becomes the reason your maintenance program got better — or just a story people tell.

Why the first 30 days matter so much

A major breakdown creates something rare in a small plant: attention. For a few weeks, the plant manager wants to talk about maintenance, operators remember exactly what they saw and heard before the failure, and nobody argues that PMs are a waste of wrench time. That window closes fast. Once production catches back up, the urgency fades, and the plant drifts back to whatever it was doing before — which is usually whatever led to the breakdown.

So treat the aftermath as a project with a deadline. Not a vague intention to “do better,” but a sequence of concrete steps with dates on them. Here's the sequence that works.

Days 1–3: Capture everything while it's fresh

Before anything else, write down what happened — in more detail than feels necessary. Memory degrades shockingly fast, and in two weeks the six people involved will remember six different versions.

  • The timeline. When was the problem first noticed? By whom? What did it look, sound, or smell like? When did the machine actually stop? When did repair work start, and what were the delays between those points?
  • What the repair took. Parts used (and where they came from — shelf stock, emergency freight, cannibalized from another machine), labor hours including overtime, outside help called in.
  • What it cost. Repair cost plus lost production. Be honest about the production number — count the missed shipments and the overtime spent catching up, not just the repair invoice. If you've never put a number on downtime before, our downtime cost calculator gives you a defensible starting point.
  • Early warnings. Ask the operators directly: was anything different in the days before? Noise, vibration, heat, leaks, error codes that got cleared and ignored? This is the question that ages worst, so ask it now.

Put all of it in one work order record, not a binder page or an email thread. You'll be referring back to it for the next month, and it becomes the anchor of this asset's history going forward.

Days 3–7: Find the actual root cause

The first explanation offered is almost never the root cause. “The bearing failed” is a description, not a cause. Why did the bearing fail? It ran dry. Why did it run dry? The grease point was missed. Why was it missed? It's behind a guard that takes fifteen minutes to remove, so it quietly fell off the routine. Now you have something you can fix.

You don't need a formal methodology — asking “why” four or five times in a room with the people who did the repair gets you most of the way. A few traps to avoid:

  • Stopping at a person. If the root cause you land on is “the tech forgot,” keep going. Why was forgetting possible? A schedule that lives in someone's head fails when that someone is on vacation. Fix the system, not the tech.
  • Stopping at the part. “Cheap import bearing” might be true, but if the same spec bearing runs three years on the identical machine next to it, the part isn't your root cause.
  • Accepting “it just wore out.” Sometimes true! But then the question becomes: did we know its expected life, and were we watching for it? Wear-out you saw coming is a planned replacement. Wear-out that took down a line for four days is an information failure.

Days 7–14: Fix the failure mode, then look sideways

First, close the loop on this machine: whatever the root cause was, put the countermeasure in place and put it on a schedule. If the grease point was inaccessible, add the remote grease line, create the PM, and set the interval.

Then ask the higher-leverage question: where else is this same failure waiting? If a missed lube point took down this gearbox, walk the plant and list every other hard-to-reach lube point. If a $40 sensor failure cascaded into a $15,000 repair, find the other places a cheap component can take out an expensive one. A single breakdown is usually a sample from a whole class of vulnerabilities, and the sideways look is how one failure buys you protection across the plant.

This is also the moment to rank what you're protecting. If you've never formally ranked your equipment, the breakdown just handed you a vivid example of what “critical” means — use that energy to run a quick ABC criticality analysis across your asset list so the next round of PM effort lands on the machines that can actually hurt you.

Turn the postmortem into a schedule

RunTight turns the countermeasures from your breakdown review into recurring PMs with checklists, and keeps the failure history attached to the asset — so the lesson survives longer than the memory. Free for teams up to 25.

Get Started Free

Days 14–30: Convert attention into a program

By week three, the machine is running and the pressure is off. This is where most plants stop — and where you shouldn't. You have maybe two more weeks where the breakdown is fresh enough to justify changes. Spend them locking in structure:

  1. Put the new PMs on a real schedule. Not a whiteboard, not a shared spreadsheet with one owner. Recurring work orders that generate themselves and show up as overdue when they're skipped. If you're starting from nothing, the small-factory PM program guide covers the full setup in an afternoon.
  2. Start logging every failure, not just the big one. The four-day breakdown got everyone's attention, but it's the pattern of small failures that predicts the next big one. Downtime tracking by cause is step one of cutting unplanned downtime, and it only works if logging starts before the next incident.
  3. Fix the spare that burned you. If the repair waited two days on a part, decide now whether that part lives on your shelf. Long-lead-time spares for critical machines are the cheapest insurance you can buy.
  4. Ask for the money while the wound is fresh. You calculated what the breakdown cost in days 1–3. A request for a stocked spare, a monitoring gauge, or a few hours a week of PM labor is an easy conversation when it's one-tenth of a number the plant manager just lived through.

The one-page breakdown review

If you formalize just one thing out of all this, make it a one-page review you fill out after every significant failure: what happened, timeline, total cost, root cause, countermeasure, sideways findings, and the date the new PM was created. Ten minutes to fill in, and it forces the loop closed — no review complete until a countermeasure exists on a schedule.

A plant that runs that loop gets measurably harder to surprise with every failure. A plant that doesn't just gets more stories. Thirty days from now, one of those will be true of yours — and it's decided this week, not eventually.

Share this article

One practical maintenance guide, every Saturday

PM schedules, checklists, and reliability tactics for small teams — the same guides published here, in your inbox. No sales pitches, unsubscribe anytime.

Ready to ditch the spreadsheet?

RunTight gives your shop automated maintenance scheduling, mobile work orders, and parts tracking — free for up to 25 users, no per-user fees.

Get Started Free