GuideAug 22, 2026·6 min read·

Run-to-Failure vs. Preventive Maintenance: When Each Actually Makes Sense

Preventive maintenance isn't always the answer. For some assets, letting them run until they break is the cheapest, smartest strategy available. The trick is knowing which assets those are — and being honest about the ones that aren't.

The dirty secret: run-to-failure is a legitimate strategy

Maintenance software companies don't love saying this, but here it is: not every asset deserves a PM schedule. Run-to-failure — deliberately letting a piece of equipment operate until it breaks, then fixing or replacing it — is a recognized, defensible maintenance strategy. For the right assets, it's the cheapest option by a wide margin.

The key word is deliberately. There's a world of difference between choosing run-to-failure for a specific asset after thinking it through, and defaulting into reactive maintenance for everything because nobody set up a program. The first is a strategy. The second is chaos with a nicer name.

Deliberate run-to-failure means you looked at the asset, decided a failure is cheap and tolerable, made sure you can recover quickly, and wrote that decision down. Reactive chaos means every failure is a surprise, every repair is an emergency, and nobody knows whether the pump that just died has failed twice this year or ten times. Same lack of PMs — completely different outcomes.

When run-to-failure makes sense

An asset is a good run-to-failure candidate when most or all of these are true:

  • It's cheap to replace. The replacement costs less than a year of PM labor and parts would.
  • It's quick to swap. A tech can change it out in under an hour with parts on hand.
  • Failure poses no safety risk. Nobody gets hurt when it dies.
  • Failure doesn't stop production. There's redundancy, or the asset simply isn't in the production path.
  • Failure is obvious. When it breaks, someone notices immediately — it doesn't fail silently and cause damage downstream.

Real examples: a general-ventilation exhaust fan with a shelf spare. A small transfer pump where a second pump sits piped in parallel. Light fixtures — nobody PMs a light fixture. Fractional-horsepower motors on non-critical equipment where the replacement costs $80 and takes twenty minutes.

Office HVAC is a useful example of the nuance. The filters get changed on a schedule — that's a PM, because a clogged filter degrades the whole system. But the small fan motor inside a rooftop unit? Many teams sensibly run it to failure, because greasing and inspecting it costs more over its life than just swapping it when it quits. One system, two strategies, both correct.

When run-to-failure is a costly mistake

The same logic that makes run-to-failure smart for an exhaust fan makes it dangerous elsewhere. Be suspicious of run-to-failure when any of these apply:

  • Single points of failure. If this asset stops, production stops. The air compressor feeding the whole plant is not a run-to-failure candidate, no matter how reliable it's been.
  • Long lead-time parts. A gearbox with a 12-week lead time turns a “cheap failure” into three months of workarounds. If you can't stock the spare, you can't afford the surprise.
  • Failures that cascade. Bearings are the classic case: a $30 bearing that runs dry doesn't just fail — it scores the $4,000 shaft it rides on, and now a grease-gun PM you skipped has become a machine rebuild.
  • Safety and compliance equipment. Fire suppression, emergency stops, guards, hoists, pressure relief valves. These aren't candidates. Inspection schedules on this equipment are often a legal requirement, not a judgment call.
  • Failure means scrapped product. If a chiller failing mid-batch ruins the batch, the real failure cost isn't the chiller repair — it's the product, the cleanup, and the missed ship date.

The math that settles it

Strip away the philosophy and the decision is one comparison:

Annual PM cost vs. failure probability × total failure cost.

Annual PM cost is the labor and parts you'd spend maintaining the asset for a year. Total failure cost is everything a failure actually costs: the repair, the downtime, any collateral damage, any scrapped product. Failure probability is roughly how likely the asset is to fail this year without PM — your failure history is the best source, a rough estimate is fine to start.

Where PM wins: a production conveyor gearbox. PM cost is about $300 a year — oil changes and an inspection. Without PM, call it a 30% chance per year of a failure that costs $2,000 in repair plus $6,000 in lost production over a down shift. Expected failure cost: 0.3 × $8,000 = $2,400 a year. Spending $300 to avoid an expected $2,400 is an easy call.

Where run-to-failure wins: a washroom exhaust fan. PM would run about $150 a year in inspection and lubrication labor. Without PM, maybe a 20% chance per year it fails — and the failure costs $200 for a new fan plus 45 minutes to swap it, with zero production impact. Expected failure cost: 0.2 × $250 = $50 a year. Paying $150 a year to avoid $50 a year is a bad trade. Let it run, keep a spare on the shelf.

You don't need precise numbers. Even rough math makes most decisions obvious — and it exposes the assets where you've been paying for PM out of habit, or gambling on run-to-failure without realizing the stakes.

Track both strategies in one place

RunTight handles PM schedules for your critical assets and failure logging for your run-to-failure assets — so the data tells you when it's time to switch an asset from one strategy to the other. Free for teams up to 25.

Get Started Free

Deliberate run-to-failure still needs three things

Choosing run-to-failure doesn't mean choosing to do nothing. It means replacing scheduled maintenance with three cheaper safeguards:

  1. A stocked spare — or a known lead time. If the plan is “swap it when it dies,” the spare has to exist. Either it's on the shelf, or you've verified the supplier can ship one in a day or two. “We'll figure it out when it happens” is not a plan.
  2. A documented swap procedure. Even a short one: where the spare lives, what tools are needed, lockout points, and any settings to transfer. The whole point of run-to-failure is a fast, boring recovery — that only happens if the swap doesn't depend on the one tech who's done it before.
  3. Failure logging. Every failure gets recorded, even the five-minute ones. This is the safety net for the entire strategy. “Cheap and rare” is the assumption run-to-failure rests on, and the only way to notice it's become “expensive and monthly” is a log that shows the same fan has been replaced four times since March.

That third one is where informal run-to-failure quietly fails. Without a log, each failure feels like an isolated event, and an asset can bleed money for a year before anyone connects the dots.

Audit your asset list

Here's the practical takeaway: go through your asset list and make the call explicitly for each one. PM or run-to-failure. Not by gut feel in the moment, but as a recorded decision — this asset gets a schedule, that one gets a shelf spare and a swap procedure. Most small plants land somewhere around PM for the critical 30–40% of assets and run-to-failure for the rest, and that's a perfectly healthy split.

Then let the failure history vote. When a run-to-failure asset starts failing more often or more expensively than you assumed, move it onto a PM schedule. When a PM has never once caught a problem on a cheap, redundant asset, consider dropping it. The strategy isn't sacred — the data is. The teams that get this right aren't the ones that PM everything. They're the ones that know, for every asset, which strategy they chose and why.

Share this article

One practical maintenance guide, every Saturday

PM schedules, checklists, and reliability tactics for small teams — the same guides published here, in your inbox. No sales pitches, unsubscribe anytime.

Ready to ditch the spreadsheet?

RunTight gives your shop automated maintenance scheduling, mobile work orders, and parts tracking — free for up to 25 users, no per-user fees.

Get Started Free