A pilot is not a trial subscription. It’s a bounded test with a start date, an end date, and a question it’s supposed to answer: does this system tell you something true about your floor that you didn’t already know, and does that something change what you do next. Most pilots fail to answer that question cleanly, not because the software didn’t work, but because nobody wrote down what “worked” meant before the gateway went in. Thirty days is short. It’s long enough to prove real things and too short to waste on vague ones.
Write the criteria down before day one
The single biggest predictor of a pilot that goes nowhere is starting it without a written definition of success. Not a feeling, not “let’s see how it goes,” an actual sentence with a number in it: “the system correctly flags the top three downtime reasons on press 4, matched against what the supervisor already knows happened.” Or: “changeover time on line 2 is measured within five minutes of a stopwatch check, three separate times.” Write it down, share it with whoever’s running the pilot on the vendor’s side, and revisit it at the end. That last step matters more than the first two. A criterion nobody checks against at the end wasn’t a criterion. It was a hope.
The reason this has to happen before day one, not during week three, is that a moving target always moves toward whatever the system happened to be good at. If you don’t pin the goal down early, a pilot that measures OEE well but never gets the shift schedule right will end up redefining success around OEE, unnoticed, and let the shift problem slide. Decide up front what matters, including the parts you expect to be hard.
The first thing a pilot has to show: a number nobody in the room disputes
Every pilot eventually produces a dashboard. The dashboard is not the proof. The proof is what happens when you put that dashboard’s number in front of the people who actually run the machine and ask if it’s right. A pilot has done its job on this front when a supervisor looks at a downtime Pareto and says “yeah, that changeover on the Tuesday run has been a problem for months,” not when they say “huh, that doesn’t look right” or, worse, shrug because they’ve stopped paying attention to the number entirely.
This is a harder bar than it sounds. A system can be technically accurate and still produce a number nobody trusts, usually because a threshold was set wrong, a shift boundary doesn’t match how the floor actually runs, or a machine that should read as running is instead sitting on a stale value. Thirty days is enough time to catch and fix that kind of mismatch once, sometimes twice. If it’s still happening in week four, that’s data too, just not the kind you wanted.
The second thing a pilot has to show: what it does when things aren’t clean
Every floor has a machine that doesn’t fit the clean case. An old control with a thin data feed, a shift that doesn’t start on the hour, a week where the schedule fell apart because three people called in. The second thing a pilot needs to prove, separate from whether the main number is right, is what the system does when it hits one of those. Does a machine it can’t measure yet show a dash and say so plainly, or does it fill the gap with a default that looks like data without flagging that it isn’t. Does a shift with nothing to report get left off a comparison, or does it get averaged into a number that no longer means what it claims to mean.
This matters more than it seems like it should, because the failure mode here is silent. A wrong Pareto gets caught the first time someone who knows the floor looks at it. A dashboard that fabricates a number to fill a gap can look completely normal for months before anyone traces a bad decision back to it. Thirty days won’t cover every edge case your floor has, but it’s enough time to find the two or three that matter most and watch how the system handles them.
What thirty days actually buys you
A pilot scoped correctly stays small. One gateway on one line, or two gateways across a wider slice of the plant, running for a month, with the hardware included in the price and credited back if you continue past the trial. That scope is not a limitation. It is what makes the test mean something. A month is long enough to see a full production cycle, including whatever irregular week always shows up, and short enough that the vendor has to prove the case fast instead of stretching a demo out indefinitely. If a vendor’s pilot structure doesn’t include a hard end date and a defined machine count, ask why, because an open ended trial with a moving scope is a sales process wearing a pilot’s clothes.
Read-only access matters here too, and it’s worth confirming in writing before the pilot starts, not discovering by accident. A system that only reads from your controllers, and never writes back to them, has a bounded worst case if something goes wrong: a gap in a chart, not a wrong command sent to a machine mid-cycle. That’s a fact worth pinning down for any vendor’s pilot you evaluate, including the ones you pass on.
What to walk away from
Some patterns during a pilot are worth ending it early over, before you’ve sunk a full month into something that was never going to answer your question. A vendor who can’t tell you in writing what the pilot is supposed to prove, and instead talks in terms of “seeing how it feels,” hasn’t built a system designed to be measured. They’ve built one designed to be liked. A dashboard that reads suspiciously good in week one, numbers that look better than what the floor already believes, is worth more scrutiny than one that reads rough, because a system tuned to flatter you on day one usually got there by loosening a standard somewhere, not by finding real capacity.
Walk away, too, from a pilot where you can’t get your own data out. If the only way to see what the system found is logging into someone else’s dashboard, you haven’t run a pilot. You’ve been given a preview. Ask for an export, a CSV, anything that proves the numbers exist independent of the interface showing them to you, and treat a vendor who can’t produce one as an answer in itself.
Last, watch for a pilot that never produces a specific, checkable finding. A month should end with something concrete: the changeover on job 4187 that’s been running long, the shift comparison that finally accounts for the informal night crew, the machine that turned out to be down twice as often as anyone realized. If a pilot ends with only a general sense that the dashboard “looked nice,” it proved nothing, no matter how many weeks it ran.
Quick recap
- Write the success criteria down before day one, with numbers in them, and check the pilot against them at the end
- The first proof point is a number the floor recognizes as true when they see it, not a number they shrug at
- The second proof point is how the system handles a machine or shift it can’t measure cleanly, a visible gap or a fabricated one
- A well scoped pilot runs about a month, with hardware included and credited if you continue, and a fixed machine count
- Confirm read-only access in writing before the pilot starts, not after
- Walk away from a pilot with no written criteria, a suspiciously perfect week one, no data export, or no specific finding at the end
Pilot criteria FAQ
What should be written down before a pilot starts? A specific, checkable statement of what success looks like, ideally with a number in it, agreed on by both sides and revisited at the end of the trial. Without it, a pilot tends to redefine success around whatever the system happened to do well.
What are the two things a pilot needs to prove? First, that its core numbers match what the floor already knows to be true when checked against a supervisor’s own knowledge. Second, that it handles a machine or shift it can’t measure cleanly by showing a visible gap instead of a made-up number.
When should you walk away from a pilot early? When there’s no written success criteria, when week one looks suspiciously better than the floor’s own sense of things, when you can’t export your own data, or when the trial ends without one specific, checkable finding.