Concept

How to Verify a Fix Actually Held

Before and after windows done right, the confounders that fake a win, and when it's fair to call an improvement zero instead of stretching to find one.

Last reviewed September 12, 2026

“We fixed it” is a claim. It isn’t proof. Most plants act on that gap constantly without noticing: maintenance replaces a part, a supervisor tweaks a setting, someone declares the problem solved, and everyone moves on, because the machine hasn’t done the annoying thing again this week. Silence for a week is just an absence of a data point, and treating that absence as a measured result is where a lot of confident-sounding fixes actually come from.

Verifying that a fix actually held means setting up a real comparison of what happened before against what happened after, with enough rigor that the answer survives someone asking “are you sure.” It’s a habit worth building for every fix, not merely the ones that already feel important.

Set the before and after windows first, not after you like the result

The most common way a verification gets gamed, usually without anyone intending to game it, is picking the comparison windows after seeing how the data looks. A person convinced the fix worked will, without really deciding to, choose a “before” window that includes a rough patch and an “after” window that starts right after the fix and conveniently avoids the one bad day that happened three weeks later. That’s not deception exactly. It’s how confirmation bias operates when nobody’s set the rule in advance.

The fix is mechanical: decide both windows before you know the outcome. A reasonable default is a period before the change equal in length to a period after it. Both need to cover normal variation: a full week at minimum for most floor problems, and sometimes a full production cycle when the issue is tied to a specific product mix that doesn’t run every day. Lock the window lengths, then let the data fall where it falls.

Equal windows, set before the outcome is known Before and after, same length, decided in advance Before window 14 days, ending at the fix Fix applied After window 14 days, starting at the fix Window length chosen before anyone saw the after data. No cherry-picking either edge.
Set both edges before you look at the outcome. Choosing them after the fact is how a wash gets reported as a win.

The confounders that fake a win

A machine’s output moves for dozens of reasons that have nothing to do with the fix you just made, and the before-and-after comparison needs to at least ask whether one of them is really what moved the number.

A schedule change. If the after window happens to run an easier product mix, less changeover-heavy, more of the machine’s best-suited part, the number improves regardless of whether the fix did anything.

A different operator or crew. A fix implemented right before a schedule rotation can get credit that actually belongs to a more experienced crew now running the shift.

Seasonal or material effects. A material lot that happens to machine more easily, humidity affecting a process that’s sensitive to it, can shift a number the exact same week a fix went in, for reasons entirely unrelated to it.

Regression to the mean. A machine gets fixed specifically because it had an unusually bad stretch. Some of that badness was likely a temporary run of unlucky variation to begin with, and it would have partially reverted on its own even with no fix at all. Comparing against the worst stretch on record inflates the apparent size of any improvement.

A yes to any of these four questions just means the raw gap needs a harder look before anyone reports it as the fix’s number. A genuine verification asks the question up front, before reporting the result, instead of after someone challenges it in a meeting.

Comparing against the right baseline

A single machine’s own history, before and after, is usually a better baseline than comparing it to a different machine, because two machines rarely run close enough to identical conditions for the comparison to isolate the fix cleanly. When a same-machine comparison isn’t available, because the fix applies plant-wide, for instance, a comparable but unaffected machine running similar work can stand in as a rough control, showing what would have happened anyway. If that control machine also improved over the same window, some of your “win” was actually a broader trend, not the fix.

When it’s fair to call it zero

The instinct once a fix is deployed is to want it to have worked, and that instinct pushes a lot of verifications toward finding some improvement to report, even a small one, instead of admitting the numbers didn’t move. Resisting that pressure is the actual discipline here. If the after window’s numbers sit within the normal range of variation the before window already showed, day to day noise, not a real shift, the right call is no measurable change, not a strained argument for a marginal win.

Calling it zero is the verification process succeeding, not failing. A fix that gets marked “no change” tells you the actual root cause is still out there, unaddressed, which is more valuable to know now than six months from now when the same failure returns and everyone’s confused about why the “fix” didn’t stick. A tracked result that separates verified improvement from no measurable change from worse, instead of assuming any action taken counts as a win, is what keeps a fix list grounded in reality instead of a scoreboard of good intentions.

What this looks like as a built-in habit

The discipline above works whether it’s done in a spreadsheet or a dedicated tool, but a system that does it automatically removes the temptation to skip the hard parts under deadline pressure. Spall’s Verified Savings ledger is a worked example of the shape this takes end to end. Linking a fix, a completed job, a work request, a dated note, to the insight or loss it addressed opens an outcome window that has to elapse before anything gets marked. Once it does, the ledger states the result plainly: Verified, No change, Worse, or Insufficient data, a measured outcome instead of a status assumed just because an action got taken. The subtitle on the page says outright that this tracks a correlation between an action and a change, not a controlled experiment, so nobody reading it mistakes it for more certainty than the comparison can actually support.

While that window is still open, the linked item shows a Measuring badge with a line explaining when to expect the result, instead of either a premature verdict or a finding that disappears from view while it waits. That single design choice, refusing to report a result before the window closes, is most of what separates a real verification from a plant just deciding a fix worked because nobody complained about the machine this week.

Four states a fix moves through before it counts A fix, tracked from claim to result Open Acted Measuring Verified or No change, or Worse Only the last box is a result. Everything before it is a claim still waiting on its window to close.
A fix isn't a result until the window closes and the comparison actually runs. Everything before that is still a claim.

Quick recap

  • Set before and after window lengths before seeing the outcome, not after the data already looks favorable.
  • Check for a schedule change, a crew change, a seasonal or material shift, and regression to the mean before crediting the raw gap to the fix.
  • A same-machine before-and-after comparison beats comparing across two different machines whenever it’s available.
  • If the after window sits inside normal variation, the correct answer is no measurable change, not a stretched case for a small win.
  • Calling a fix “no change” is the verification process succeeding, not failing. It means the real cause is still out there.

Related guides

From the Help Center

$200/machine/mo · pilots from $4,500 · hardware included

See full pricing →

Bring your machine list. We'll tell you exactly what plugs in.

30 days, hardware included, line pilot $4,500 or plant pilot $8,500, fully credited when you expand.