“There’s a 62% chance this press fails in the next week” is a strong claim for a maintenance planner to act on. It’s worth asking where the 62% came from before scheduling around it. Spall’s failure-risk board is built so that question always has an answer: the count of past failures on that exact machine, and the math run over them, both sitting right next to the number.
What counts as a failure, and what doesn’t
Not every stop on a machine’s history is a failure a planner would recognize. A quality check, a short jam, an unlabeled gap, none of those pages anyone. Spall counts a downtime event as a real failure only when it clears one of two bars: its reason code is breakdown class, or a repair-class work item was actually raised against it. Everything else, coded or not, stays out of the count.
Two cleanup steps run before anything gets counted. Overlapping or duplicate spans on the same machine, two logging paths catching the same physical stop, collapse into one episode instead of two, so a single breakdown never gets counted twice. And a merged episode shorter than five minutes gets dropped entirely, the same minor-stop cutoff OEE math already uses elsewhere in Spall, because a few minutes of flicker isn’t a breakdown a technician would recognize as one.
What’s left after both steps is a machine’s real failure history: repair-grade stops, deduplicated, timed accurately. That’s the only material the estimate below is built from.
MTBF, and whether it’s getting worse
Once there’s a real failure history, the gaps between consecutive failures become the raw material for two numbers. Mean time between failures, MTBF, is the average of those gaps. A second average, taken over just the newer half of the history, gets compared against the overall one: if the recent gap has shrunk to three quarters of the long-run average or less, the machine is flagged degrading. If it’s grown to a quarter more, it’s flagged improving. Otherwise it’s stable. A machine is only ever compared against its own past, never against a fleet average or another customer’s equipment.
The odds of a failure this week
This is the part that actually predicts something, so it’s worth being precise about how. Spall does not fit a curve to a machine’s failure pattern and extrapolate from it. It counts. Take every historical gap between failures that ran longer than the time since this machine’s last failure, those are the only fair comparisons, a gap that ended after 3 hours can’t tell you anything about a machine that’s already been running for 3 days. Of those longer historical gaps, the fraction that ended within the next seven days becomes the probability shown on the board.
When a machine has already run longer than every historical gap on record, there’s nothing left to compare it against, no past run ever survived this long, so there’s no fraction to compute. Instead of stretching the model past its own data, the board marks the machine overdue and stops there, with a line telling you it’s already past its longest observed run without failing and is worth an inspection.
What a wrong guess is worth
A probability by itself doesn’t tell a planner whether to act now or next month. Spall multiplies the probability by the average repair time this machine’s own past failures took, then by its resolved dollar-per-hour rate, to get a risk-weighted exposure figure. It’s labeled that way for a reason: an expected value across a range of outcomes, not a bill anyone’s actually going to pay. A 20% chance of an eight-hour repair on a $600-an-hour line prices out the same as a 40% chance of a four-hour one, and the board’s number reflects that, even though the two situations call for different responses.
When it says nothing
Below five real failure-to-repair cycles, the board doesn’t produce a percentage at all. It states the count against the floor directly, something like “not enough failure history to estimate risk yet,” and moves the machine to a list of assets still building history instead of a fleet-wide estimate stretched over almost no data. A second, independent check catches a subtler case: even with five or more cycles, if the machine’s own average gap between failures is itself close to that five-minute minor-stop cutoff, the failures are happening too close together to call a credible cadence, and the board abstains there too rather than reporting a number built on noise.
A number’s reliability also grows with how much history backs it. The board carries a confidence value that climbs toward its maximum around sixteen real intervals, so a machine sitting right at the five-cycle floor shows a lower confidence than one with years of failures behind it, even if the raw percentage looks identical.
A vendor selling a black-box health score can’t answer “how many failures is this built on” because the model was trained somewhere else, on other people’s equipment, and tuned to look confident regardless of how much this particular machine has actually failed. Every number on Spall’s board answers that question by design, the count is printed right next to the percentage.
Where to see it
The Maintenance page’s Failure risk board shows every asset that’s cleared both floors, sorted overdue first, then by probability, then by dollars at risk, with the machine’s MTBF, its trend, and its exact failure-cycle count printed alongside each entry. Assets still below the floor sit underneath it, named plainly with the reason. MTTR and MTBF for a Plant Floor covers the two averages this board is built from in more depth, and Downtime Reason Codes covers the coding discipline that decides which stops even become candidates for this count in the first place.
Quick recap
- A failure only counts when its reason is breakdown class or a repair-class work item was raised against it, deduplicated and floored at five minutes so flicker and double-logged stops don’t inflate the count.
- MTBF and a degrading/improving trend compare a machine’s recent failure gaps against its own longer history, never a fleet average.
- The seven-day probability is a direct count: the fraction of historical gaps longer than the current run that ended within a week, no curve fit.
- A machine running longer than every observed gap gets marked overdue instead of a stretched percentage past the edge of its own data.
- Below five real failure cycles, or with failures too close together to be a credible cadence, the board states the reason instead of a number, and confidence grows with how much real history backs the ones it does show.