The pilot was the easy part
Ten users, curated inputs, someone watching. Then it meets a Tuesday.
FORTY CASES, THEN TEN THOUSAND
The exceptions were always there. Only volume makes them visible.
A pilot works partly because of what has been removed from it: curated inputs, volunteer users, someone watching, and low volume. Go-live restores all four on the same day. The model is unchanged; the conditions are not, which is why most pilots never become systems.
The demo worked. Everyone in the room saw it work. That is usually the last uncomplicated moment in the project.
What happens next is well enough documented to have stopped being a surprise. RAND’s 2025 analysis found that roughly eight in ten AI projects fail to deliver the business value they were approved on. IDC’s figure is starker at the pilot stage: of thirty-three proofs of concept, four graduate. MIT Sloan, looking specifically at generative AI, put the share of pilots that fail to scale at ninety-five per cent.
The numbers differ because the definitions differ — abandoned before production, completed but underdelivering, live but never measured. The direction does not differ. Most pilots do not become systems.
The reflex is to treat this as a technology problem, and it almost never is. The same model, the same prompt, the same integration that worked on Thursday afternoon in front of eight people fails on the following Tuesday in front of nobody. Nothing about the model changed. Everything about the conditions did.
What a pilot quietly removes
A pilot is not a small version of production. It is a different thing wearing the same clothes, and it works partly because of what has been taken out of it.
The inputs are curated. Someone chose the test cases. Not maliciously — sensibly, because you cannot demonstrate anything with a hundred thousand rows of whatever happened to arrive. But the cases that were chosen were the ones that made sense, and the ones that make sense are not representative of what arrives on a Tuesday.
The users are volunteers. Pilot users want it to work. They rephrase when the answer is odd. They know what the system is for. They forgive. The people who arrive after go-live did not volunteer, do not know what it is for, and will type whatever they were going to type anyway.
Someone is watching. During a pilot there is a person who notices when something looks wrong. That person is not staffed after go-live, and in most organisations nobody notices that they have gone.
The volume is small. Problems that appear once in five hundred cases do not appear at all in a pilot of forty. They appear reliably at ten thousand.
Take all four away at once — which is exactly what go-live does — and you are not scaling the pilot. You are running an experiment nobody designed.
The three failures that look like model failures
Input drift. The distribution of what arrives changes, and nothing announces it. A new form field, a partner sending a different file layout, a seasonal shift in what customers ask about. The model handles it the way models do: it produces an answer with the same confidence as always. This is the failure most often misdiagnosed as the model degrading. The model is unchanged; its diet is not.
The long tail arriving all at once. Every operation has a tail of exceptions. In a pilot they are noise; in production they are a queue. Automation that covers eighty per cent of cases is not eighty per cent of the work, because the twenty per cent was where the hours always were.
No feedback path. In a pilot, wrong output is discussed in the room. In production, wrong output is worked around silently. The person who receives it fixes it by hand, does not report it, and the system’s owner sees a dashboard that says everything is fine. The failure is real and invisible at the same time. This is the most expensive of the three, because it can run for months.
Why measurement is not a phase
The pattern in the research is consistent on this point. MIT’s work identified measurement built in from the start as the thing separating the small group that scaled from the rest. The organisations that scaled were not the ones with better models. They were the ones that could tell.
The instinct is to treat measurement as something added once the thing works. That ordering does not survive contact with production, because by then the baseline is gone. You cannot show that a system saved four hours a week if nobody recorded what the four hours were before it arrived.
What has to exist before go-live, not after:
- A baseline. What did this cost in time, errors or headcount last month? Recorded, not estimated afterwards.
- A sample somebody reads. Not a dashboard — a fixed number of real outputs, read by a person, on a schedule. Twenty a week is enough to notice drift.
- A definition of wrong. Written down before launch, so that the argument about whether an output was acceptable does not happen in the moment.
- A named owner. Not a team. A person, whose job includes this, and who is still there in month six.
None of that is technical. All of it is the difference between a system that is running and a system somebody is running.
The shape this usually takes
The version of this that costs the most is not dramatic. It looks like a system that has been live for several months and is generally believed to be working.
Someone eventually checks — because a customer complains, or a number does not reconcile, or a new person asks a question nobody had asked. The output turns out to have been wrong in a narrow, consistent way for most of that time. Not wrong enough to stop anyone, which is precisely why it ran so long. The people receiving it had been correcting it by hand and had stopped thinking of that as a correction.
Two things are usually true at that point. Nobody had read the actual output since the pilot. And when the question “who owns this” is asked, the honest answer is that the person who built it moved on and nothing replaced them.
Neither is a technology failure. Both are the reason the system was allowed to be wrong quietly.
What the ones that survive have in common
Not model choice. Not the framework. Not the budget.
They went live smaller than they could have. A narrow slice, fully instrumented, with a named owner and a real baseline — then widened once there was evidence it held. The wide launch feels faster and is the reason the twelve per cent figure exists.
They also, without exception, had somebody whose job it was to look. The systems that fail quietly fail because nobody’s calendar has an appointment with them.
Figures drawn from published research by RAND, IDC and MIT Sloan, cited inline with years. This is a reading of that work alongside patterns seen in deployment, not a study of our own.
Limits
Different sources, different definitions
The figures come from RAND, IDC and MIT Sloan, and they count different things — abandonment, non-scaling, absent ROI. They agree on direction and disagree on magnitude. Treat the range as the finding, not any single number.
This is about deployments, not models
Nothing here says anything about which model to use. The failures described are organisational and occur regardless of what is underneath.
Your operation may be the exception
Small, high-volume, low-variance workloads behave differently from the ones this describes. The fastest way to know is a conversation, not a longer article.
Questions
What percentage of AI pilots reach production?
Estimates range from about twelve per cent (IDC) to under a third, depending on what counts as production. RAND’s 2025 analysis found roughly eighty per cent of AI projects fail to deliver the business value they were approved on.
Why do pilots succeed and production fail?
A pilot removes four things: curated inputs, volunteer users, someone watching, and low volume. Go-live restores all four at once. The model is unchanged; the conditions are not.
Is this a model quality problem?
Rarely. Across the cited research the recurring causes are data readiness, integration, ownership and change management — not model capability.
What should be in place before go-live?
A recorded baseline, a fixed sample somebody reads on a schedule, a written definition of what counts as wrong, and a named owner who is still there in month six.
How do you tell whether a live system is quietly wrong?
Read its output. A fixed number of real cases, weekly, by a person. Dashboards report volume; they do not report correctness.
How long before problems appear?
Input drift and the long tail usually surface in the first two to three months of real volume — which is also when the pilot team has typically moved on.
XONIK, “The pilot was the easy part”, Field Report, September 2026.