Concept

MTTR and MTBF for a Plant Floor, in Plant Words

What mean time to repair and mean time between failures actually measure, how much history makes them trustworthy, and what to do with a number built on too little of it.

Last reviewed September 12, 2026

MTTR and MTBF show up in almost every maintenance conversation, usually spoken like everyone already agrees what they mean. Most of the time nobody’s checked. A tech means one thing by “MTTR,” a plant manager building a board deck means something close but not identical, and a quality auditor asking for the number means something built entirely differently. The definitions are simple. The disagreements are almost always about what counts as an event in the first place, not about the math.

Mean time to repair, in plant words

MTTR is how long a stop takes to close, on average, once it’s flagged as something maintenance needs to work, not from the moment the machine went down, but from the moment a repair actually starts, through to the moment it’s marked fixed. A real repair queue moves a stop through stages, someone acknowledges it, starts working it, then resolves it with a note on what actually fixed it, and each stage change gets its own timestamp. MTTR is the average length of that whole path across a set of repairs, not a guess made at the end of the week from memory.

The plant-floor version of this number is more useful than the finance-office version precisely because it’s built from real timestamps, not from a technician filling in “about two hours” on a form after the fact. A technician under pressure to look good rounds down. A system timestamping acknowledge, start, and resolve automatically doesn’t have an opinion about how the number looks.

Where MTTR is measured on a repair's timeline One repair, marked on its own timeline Stop Acked, 5m Repair starts, 12m Resolved, 40m later This span is MTTR
Every stage carries its own timestamp. The repair span, not the whole outage, is what MTTR averages.

A resolve step that requires an actual note on what fixed it, not merely a tap, is what keeps MTTR trustworthy over time. A repair marked done with no note is a number with no explanation behind it, and a plant that’s serious about lowering MTTR needs those notes to find the pattern, the same bearing failing on the same machine, the same part that’s always slow to source.

Mean time between failures, in plant words

MTBF is the average stretch of running time between one genuine failure and the next, on a given asset. The word doing the real work in that sentence is “genuine.” MTBF counts real breakdowns, an actual failure or a stop that had a repair opened against it, not every minor jam or short stop a machine has. A machine that pauses for a sensor re-trigger forty times a shift hasn’t failed forty times, and folding those into an MTBF calculation would make every asset’s reliability number meaningless, drowned in events that were never failures in the first place.

That distinction is also what makes MTBF worth tracking as an early signal instead of just a scorecard entry. A machine whose MTBF is visibly trending down, its real failures happening closer together than they used to, is telling you something is drifting before it fully breaks. That’s a real, useful signal, and it costs nothing extra to generate: it’s built from the same repair records MTTR already comes from, no new sensor, no new program, just consistent logging of what counts as a real failure.

MTBF trending down across four real failures One machine, four real failures, four months 40 days 38 days 22 days Gaps shrinking: 40, then 38, then 22 days between real failures.
Three shrinking gaps in a row is a trend worth a maintenance look, not a coincidence to wait out.

MTTR is not the same clock as time to acknowledge

The two numbers get confused constantly, and the confusion causes real damage. Time to acknowledge is how long a call sat before anyone responded. MTTR starts once a repair is actually being worked and ends when it’s fixed. A five minute acknowledge time on a repair that then takes four hours to complete is still a five minute acknowledge time, and it’s a perfectly good one, the tech got there fast. Folding that four hour repair into a response-time number makes a fast-responding team look slow, and folding a slow acknowledge into MTTR makes a fast-repairing team look slow too. Andon Response Time: How Fast Is Fast Enough covers the acknowledge side of this in depth. Keep the two clocks separate and neither can hide behind the other.

Using both together to spot a repeat offender

MTTR and MTBF earn their keep together more than either does alone. An asset with a fine MTTR, every individual repair closes fast, can still be bleeding the plant dry if its MTBF keeps shrinking, because a five minute fix repeated every three days costs more total downtime and more total interruption than a ninety minute fix that happens twice a year. Watching both side by side catches that pattern: an asset showing up again and again in the repair queue is where an hour of root-cause work pays back fastest, even when no single repair on it ever looked alarming by itself.

What sample size makes them trustworthy

Both numbers are averages, and an average built on two data points is really a coincidence dressed up as a statistic. One repair that happened to take ninety minutes only establishes that one repair took ninety minutes. Calling that a trend is how a maintenance team ends up chasing a number that was never real.

There’s no single magic count that turns a handful of repairs into a trustworthy average, but a rough rule holds up in practice: treat anything under five or six events as too thin to act on with confidence, and trust the number once you’re past a dozen. A newly connected machine, or one whose failures are rare, one every few months, might take a long stretch of calendar time to reach that count, and that’s fine. A reliability number that shows a dash instead of a computed figure until there’s enough history is worth more than one that fills the gap with a guess dressed up as precision.

What to do with a bad reading

A bad MTTR or MTBF reading, whether that means suddenly worse or just alarmingly good, deserves a second look before it drives a decision. A single unusually long repair can drag an average up for weeks on a machine that otherwise repairs fast, and a single unusually short gap between failures can make MTBF look like it’s collapsing when it’s really one bad week. Look at the individual events behind the average before reacting to the average itself. Was one repair long because parts had to be sourced from three states away, a real and repeatable problem, or because the tech got pulled to a different fire mid-repair, a scheduling problem that has nothing to do with the asset.

Worth watching more than either number in isolation is the shape of the trend across several periods. One bad MTTR week is noise. Three months of MTTR creeping up on the same asset is a real signal that something about the repair process, or the failure mode itself, has changed and is worth a deliberate look before it becomes the new normal nobody questions anymore.

Quick recap

  • MTTR measures the repair span, acknowledge through resolve, not the whole outage from the moment the machine stopped.
  • MTBF counts genuine failures only, an actual breakdown or a stop with a repair opened against it, never every minor stop.
  • Both numbers come from the same repair-logging discipline, no separate program needed to start generating them.
  • Treat anything under five or six events as too thin to trust. A dash beats a guess when there isn’t enough history yet.
  • Check the individual events behind a surprising average before reacting. One outlier can move the number without meaning anything changed.
  • A trend across several periods, not one bad week, is what actually deserves a maintenance response.

Related guides

From the Help Center

$200/machine/mo · pilots from $4,500 · hardware included

See full pricing →

Bring your machine list. We'll tell you exactly what plugs in.

30 days, hardware included, line pilot $4,500 or plant pilot $8,500, fully credited when you expand.