The short answer
A pilot succeeds and the production rollout disappoints. The technology did not change, the team did not change, and the result did.
The usual explanation is scale: more volume, more edge cases, more integration. That is part of it and it understates the problem, because the more important difference is that a pilot runs under conditions chosen to make it work.
Those conditions are rarely written down. Participants volunteered. The cases were a clean subset. The build team was one message away. Leadership was watching, which made the workflow change happen. A project member quietly handled whatever the system could not.
Every one of those is removed at scale, and because none was named, nobody notices they were removed. The production deployment is measured against a pilot result produced under circumstances that no longer exist.
Naming the conditions in advance is what makes a pilot predictive. A pilot that removes them deliberately produces a lower number and a usable forecast.
Key takeaways
- Pilot conditions are chosen to make the pilot work, and they are usually not documented.
- Seven conditions account for most of the gap, and all seven disappear at scale.
- Self-selected participants are the single largest distortion, since enthusiasm substitutes for usability.
- Manual gap-filling by the project team is the most invisible, because it feels like support rather than compensation.
- A pilot designed to test what will break produces a worse number and a better decision.
- Scale-up criteria agreed before the pilot prevent the result from being reinterpreted afterward.
The seven conditions
1. Participants volunteered
Pilot groups are assembled from people who are interested. They tolerate friction, report issues constructively and work around gaps.
At scale the population is everyone, including people who did not ask for this, are measured on throughput, and will route around anything that slows them down on a busy day.
Enthusiasm in a pilot substitutes for usability. When the enthusiasm is removed, the usability problems that were being absorbed become visible all at once.
2. The case mix was clean
Pilots run on a selected subset: a defined period, a team, a product, a market. The selection is practical and it systematically excludes the complicated cases, because complicated cases are harder to set up.
Production receives the full distribution, including the exception types that make the process expensive. A system scoped against the pilot subset meets them for the first time in production.
3. The build team was accessible
During a pilot, a question gets answered in minutes. A configuration issue is fixed the same day. A confusing output is explained by the person who built it.
At scale there is a support queue, a release cycle and documentation. The same issue that took four minutes to resolve in the pilot takes four days, and user tolerance drops accordingly.
4. Leadership was paying attention
A pilot has visibility. Managers reinforce it, people prioritize it, and the workflow change that the system requires actually happens because someone senior is asking about it weekly.
In production the attention disperses across the whole population. The workflow change has to survive on its own merits, without the reinforcement that made it happen the first time.
5. The process variant was narrow
One team in one market runs one version of the process. The system was configured for that version, and it fit.
At scale the same nominal process has variants by market, business unit and customer segment, each with requirements the pilot configuration never encountered.
6. Someone filled the gaps manually
This is the most invisible condition and frequently the largest.
During a pilot, a project team member notices a case the system handled poorly and corrects it. They re-run a failed job, fix a mapping, clean an input. It feels like support and it functions as compensation.
The pilot metrics include that compensation and attribute none of it. At scale there is no project member available to do it, and the gap it was covering becomes a visible failure rate.
7. Errors were tolerated
Everyone knew it was a pilot. A wrong output was treated as a finding rather than as an incident.
In production the same output reaches a customer, a regulator or a financial record, and the tolerance that made the pilot feel smooth is gone.
Pilot and production
What changes between the two
| Condition | In the pilot | At scale | Effect when removed |
|---|---|---|---|
| Participants | Volunteers | Everyone | Usability problems surface at once |
| Case mix | Clean subset | Full distribution | Exception failures appear |
| Support | Build team directly | Queue and release cycle | Tolerance drops |
| Attention | Leadership watching | Dispersed | Workflow change stops happening |
| Process variant | One version | Many | Configuration gaps |
| Gap-filling | Project team compensates | Nobody | Failure rate becomes visible |
| Error tolerance | High, it is a pilot | Low, it is live | Incidents rather than findings |
A pilot result is a measurement taken with all seven conditions present. Production is the same system with all seven removed.
How to design a pilot that predicts
Name the conditions before you start
Write them down as a list with the production equivalent beside each. The act of writing it changes the design, because several of them turn out to be removable during the pilot itself.
Include non-volunteers
Add a group that was assigned rather than one that asked. The difference between the two populations is the most informative result the pilot produces, and it predicts the adoption curve more reliably than aggregate satisfaction.
Include the exception cases deliberately
Select the subset to include the case types you expect to be difficult rather than excluding them. A pilot that avoids the hard cases measures the part of the process that was already working.
Instrument the gap-filling
Ask the project team to log every manual intervention: what they corrected, how long it took, and what would have happened without them. That log is the single best predictor of the production failure rate, and it usually does not exist because the interventions feel like helping.
Run at least part of it without privileged support
For a defined period, route questions through the support path that production will use. The resulting friction is real information.
Test more than one process variant
If the process differs by market or business unit, include at least two versions. Configuration gaps that appear with two will appear with twelve.
Agree scale-up criteria in advance
What result justifies proceeding, what result justifies a redesign, and what result stops it. Written before the pilot, these criteria are a decision framework. Written afterward, they are a negotiation.
Reading the pilot result
| Observation | Likely production outcome |
|---|---|
| Volunteers satisfied, assigned group neutral | Adoption will require sustained manager reinforcement |
| Gap-filling log is substantial | The production failure rate will be close to that volume |
| One variant tested | Configuration work remains, scope it before committing a date |
| Exception cases excluded | The scope covers the standard path, which was already working |
| Support was privileged throughout | Tolerance in production will be lower than observed |
| Metric moved only during leadership attention | The workflow change did not take hold |
A pilot that scores well on every row is genuinely ready. A pilot that scores well because the conditions were favourable and unexamined produces a confident decision with a poor basis.
What to do when the conditions cannot be removed
Sometimes a pilot has to run under favourable conditions, because the alternative delays it past the point of usefulness.
That is acceptable if the result is interpreted accordingly. Three adjustments.
Discount the result explicitly. State in the write-up that the measurement was taken with named conditions present, and that the production figure will be lower.
Budget for the gap-filling. If a project member spent six hours a week compensating, production needs either an equivalent operational capacity or a fix for what they were compensating for.
Stage the rollout against the conditions. Expand first to a population most similar to the pilot group, then to the one least similar, and measure the difference. The gap between the two is the real adoption curve.
Where Horizon fits
Horizon is an AI-powered continuous discovery platform. Its relevance to the pilot-to-production gap is upstream: several of the seven conditions exist because the process was not understood well enough to select a representative pilot scope.
Discovery Cycles establish the case distribution, the variants by market and team, and the exception types that carry the volume, from the people who handle them. That is what allows a pilot scope to include the difficult cases deliberately rather than excluding them by accident.
The Process Library records the process variants that exist, which is what turns "test more than one variant" from a principle into a specific list. The Insights Dashboard quantifies where effort concentrates, which determines whether a pilot is being run on a part of the process that matters.
A pilot scoped from the documented process tends to select the clean subset without anyone deciding to, because the documented process is the clean subset.
Pilot design checklist
- Have you written down the conditions your pilot will run under?
- Does the participant group include people who were assigned rather than volunteered?
- Does the case selection include the exception types you expect to be difficult?
- Will the project team log every manual intervention?
- Will any part of the pilot run through the production support path?
- Are at least two process variants included?
- Are scale-up criteria agreed in writing before the start?
- Is there a baseline for the business metric, captured before the pilot?
- Do you know what proportion of total volume the pilot scope represents?
- What result would cause you to redesign rather than proceed?
- Who owns the workflow change after leadership attention disperses?
Question 4 is the one that changes pilot results most often. Teams are consistently surprised by the volume of compensation that was happening invisibly.
Common mistakes
Treating the pilot result as a forecast. It is a measurement taken under conditions production does not have.
Excluding difficult cases to get a clean run. Measures the part of the process that was already working.
Staffing the pilot with volunteers only. Enthusiasm substitutes for usability and hides the problems that determine adoption.
Leaving gap-filling unlogged. The best available predictor of production failure rate goes unrecorded.
Agreeing scale-up criteria afterward. Turns a decision framework into a negotiation that the sponsor usually wins.
Testing one variant. Configuration gaps that appear with two will appear with twelve, and the discovery happens after the commitment.
FAQ
Why do successful AI pilots fail in production?
Because pilots run under conditions selected to make them work: volunteer participants, a clean case mix, direct access to the build team, leadership attention, a single process variant, manual gap-filling by the project team, and tolerance for errors. All seven disappear at scale, and because they are rarely documented, nobody accounts for their removal.
What is the pilot to production gap?
The difference between the result measured in a pilot and the result achieved at scale, caused mainly by the removal of favourable conditions rather than by technical scaling problems. The gap is predictable if the conditions are named in advance and measured.
How do you design a pilot that predicts production performance?
Name the conditions and their production equivalents, include participants who were assigned rather than volunteered, deliberately include difficult case types, log every manual intervention by the project team, route some questions through the production support path, test more than one process variant, and agree scale-up criteria in writing before starting.
What is the most common hidden factor in pilot results?
Manual gap-filling. A project team member notices a case the system handled poorly and corrects it, which feels like support and functions as compensation. The pilot metrics include the corrected outcome and attribute none of the effort, so the production failure rate arrives as a surprise.
Should pilots use volunteer participants?
Volunteers are useful for finding issues and misleading for predicting adoption, since their tolerance for friction masks usability problems. Including an assigned group alongside them produces the comparison that matters, and the difference between the two groups is the best available indicator of the adoption curve.
How do you decide whether to scale a pilot?
Against criteria agreed in writing before the pilot started: what result justifies proceeding, what result justifies a redesign, and what result stops it. Criteria written afterward become a negotiation, and the sponsor's prior commitment tends to determine the outcome.
The pilot measured a different situation
A successful pilot establishes that the technology works in your environment under favourable conditions. That is worth knowing and it is a narrower claim than it appears.
The conditions are the variable. Name them, remove the ones you can, measure what the rest were worth, and the pilot starts predicting something.
See it. Fix it. Scale it.