What a proof of concept actually proves.
A demonstration establishes that the task is possible. Almost every organisation reads it as establishing that the deployment will work.
The pilot was tested on the easy ones.
Not through carelessness — the difficult cases are difficult to assemble, which is exactly why they get left out.
A proof of concept demonstrates that a task is technically possible on the cases it was given. It does not establish that the cases were representative, that the exceptions are handled, or that anyone will act on the output. Those three questions carry most of the remaining cost, and a successful demonstration tells you nothing about any of them.
Somebody assembles a set of examples. They pick ones where the answer is knowable, the inputs are complete, and the situation is typical. That is the right way to build a demonstration — you cannot demonstrate anything on cases nobody can adjudicate.
The examples are therefore clean by construction. Not selected to flatter; selected to be assessable. The two produce the same set.
Production sends everything. The submission with a missing field. The customer whose situation does not fit any category. The record entered by somebody who left in 2023 and used the notes field for something nobody else does.
Those are a small fraction of volume and most of the operational cost. They are also where a wrong output has a consequence.
The number worth asking for
Not accuracy on the demonstration set. The proportion of real monthly volume that the demonstration set represents.
In most pilots we have reviewed it is under two per cent, and nobody had calculated it before being asked. A 94% accuracy figure on two per cent of volume, chosen for being assessable, is a claim about a corner of the problem presented as a claim about the problem.
Engagement observation rather than a counted study. The “under two per cent” figure is drawn from pilots where the calculation was made, which is a small number and not a sample. Treat it as an order of magnitude and check your own — the calculation takes ten minutes.
Three questions a demonstration leaves open
None of them technical. A demonstration can answer the first with more work. The other two it cannot answer at all.
Were the cases representative?
Answerable, and rarely answered. Take a month of real volume, sample fifty at random rather than by selection, and run those. The number that comes back is the one that matters, and it will be lower.
What happens to the ones it cannot do?
Every system has a case it should refuse. Where those go, who looks at them, and how long they may wait is a design question — and a demonstration never has to answer it, because a demonstration has no queue.
Who will act on the output?
The demonstration is watched by people who are interested. Production sends outputs to people who are busy, and whether they act depends on whether it is their job to. That is a question about authority rather than about quality, and no amount of model work touches it. More on that in what an agent needs.
Two changes that cost a day each
Both make the demonstration less impressive and considerably more informative.
- Sample at random from real volume. A month of cases, fifty pulled without selection. The accuracy figure will drop, and it will be the true one.
- Include the ones nobody can adjudicate. Where the right answer is genuinely unclear, that is a finding about your process rather than about the system.
- Count what the set represents. As a share of monthly volume. Under five per cent means you have a demonstration rather than a test.
- Show the refusal path. What happens to a case the system cannot handle. If there is no path, the demonstration has skipped the expensive part.
- Have somebody accept one output for real. One case, actioned in the live process, by the person who would actually do it. That single step surfaces more than the rest of the pilot combined.
- Name what would make you stop. Before running it. A demonstration with no failure condition cannot fail, which means it cannot inform anything.
Limits
Not a measured finding
This is engagement observation, not a study with a sample. The figures are orders of magnitude from cases where the calculation happened to be made.
It describes consequential operations
Where a wrong output costs nothing, a clean demonstration may be perfectly adequate and the exceptions may not matter. This is not about those.
Some demonstrations should be clean
If the question is whether a technique is viable at all, a chosen set is the right instrument. The error is in what gets concluded afterwards, not in how it was built.
Questions
Are you saying pilots are a waste?
No. A pilot is the right way to establish whether something is possible. The waste is in treating "possible" as "will work", which is a reading error rather than a problem with the pilot itself.
How large should the test set be?
Large enough to contain the exceptions, which usually means fifty to a hundred cases sampled at random from a real month. Size matters less than how they were chosen.
Our accuracy dropped when we did this. Now what?
You have a real number instead of a flattering one, which is progress. The next question is whether the errors cluster — if they concentrate in one case type, that type may simply be out of scope.
What if we cannot get real data for a pilot?
Then say so in the write-up and treat the result as a capability check rather than a readiness one. The error is not the constraint; it is the conclusion drawn despite it.
Who should choose the test cases?
Not the person building the system, and not the sponsor. Somebody who runs the process daily will pick different examples, and the ones they pick are the ones that matter.
Is one real acceptance really that useful?
It is the single most informative thing in a pilot. It surfaces who signs, what they need, what happens when they decline, and whether the output arrives in a form anyone can act on.
Govil, A. (2026). What a proof of concept actually proves. The Field Report, XONIK.