Research

Where deployments stop

Sixty-one AI deployments in regulated operations, followed from demonstration to production or abandonment.

Twenty-six stopped at the same point, and it was not a technical one.

Run the readiness assessment

The finding

Of sixty-one AI deployments followed from demonstration to production or abandonment, seventeen reached production. The largest single point of loss was not technical. Twenty-six stopped at acceptance — the point where a named person had to act on an output and answer for it. Where that accountability was designed before the build, 83% reached production. Where it was deferred, 18% did.

LengthEight exhibits, full method and appendices
Basis61 deployments, 2024–2026
AccessFree, in full

What is in it

The argument, in full.

Not a preview. What the study looked at, what it found, and what separated the outcomes.

Sixty-one AI deployments were followed from an accepted demonstration through to production or abandonment, tracked from January 2024 to June 2026 in regulated operations — healthcare, banking, professional services and education — where a wrong output carries a named consequence for somebody. Thirty-four came from direct engagement with the organisations involved; twenty-seven were reconstructed after the fact from project records and interviews with the people who ran them. Seventeen reached production.

The study set out to find where deployments actually stop, rather than assuming the answer was technical. It wasn't. Seven of the sixty-one failed at capability — the system genuinely couldn't do the task well enough — and five failed at evidence, unable to answer a regulator's or auditor's question about a specific past decision. Together those account for twelve failures. Acceptance, the point where a named person has to act on a system's output and answer for it, accounted for twenty-six — more than every other stage combined. Nothing was being built at that point and nobody had scheduled it; it simply had no owner in the project plan, which the report treats as the actual finding rather than an oversight to note in passing.

The sixty-one split cleanly on one variable: whether the accountability question — who accepts the output, and what happens when they decline it — was settled before the build began or left for later. Twenty-three organisations settled it first; nineteen of those reached production, an 83% rate. Thirty-eight deferred it; seven reached production, 18%. The report is careful about what this association can and cannot support: it correlates strongly and holds within every subgroup coded, including budget, model choice, team size and vendor, but it does not establish that early accountability causes production, since organisations that settle it first may simply be better run in ways the study didn't measure.

Looking inside the twenty-six acceptance failures, four absences recur: no allocated time for the review, no default action when nothing happened, no defined consequence for declining, and no measurement of whether outputs were actually being acted on rather than just produced. The report argues these failures aren't a training problem, even though eleven of the twenty-six ran additional training after stalling and none of it resolved anything — because declining to accept an output is genuinely the only one of three available responses that carries no personal risk for the person being asked. Where a role rather than a named individual held acceptance, or where a committee held it, deployments were far less likely to ship; where a named individual held it and could also decline without a defined penalty, most still shipped.

The report's own limits are stated in advance of the analysis: sixty-one is not a representative sample, coding wasn't independently blinded, the outcome measure — reached production or not, at eighteen months — says nothing about whether the deployments that shipped were worth shipping, and the finding is scoped specifically to settings where a wrong output has a named consequence. What a reader takes away is narrower and more actionable than a general theory of AI adoption: name who accepts an output, state what they can do without asking again, and settle both before the build starts, because doing it after costs far more than doing it first.

The exhibits

Eight charts, all in the PDF.

01

Where sixty-one deployments stopped

02

Production rate by when accountability was settled

03

Outcome by when the accountability question was first raised

04

Production rate by sector

05

Production rate by who was named as accepting owner

06

Production rate by when the evidence requirement was specified

07

Months from demonstration to production, or to last activity

08

Production rate by whether the old process had a retirement date

Method

How the finding was reached.

Period January 2024 to June 2026
Population 61 deployments that reached an accepted demonstration
Setting Regulated operations where a wrong output has a named consequence — healthcare, banking, professional services, education
Basis Direct engagement (34) and post-hoc review with access to records and staff (27). Not survey.
Selection Not random. These are deployments we saw. This biases the set toward organisations that sought outside help, and toward those willing to let an outsider examine a failure.
Unit One deployment. A programme running four workflows counts as four.
Outcome Binary: in production and used, or not, at 18 months from demonstration.

Limits

What this does not show.

Written before the analysis was run, which is the only point at which it can be written honestly. Afterwards, limits come out shaped to protect what was found.

THE SAMPLE IS NOT REPRESENTATIVE

These are deployments we saw. Organisations that sought outside help are over-represented, and so are those willing to let an outsider examine something that went wrong. Both plausibly correlate with the outcome being measured.

IT DOES NOT ESTABLISH CAUSE

The cohort difference is an association. Organisations that settle accountability first may be better run in ways we did not measure. The mechanism is plausible and the correlation is strong, and neither of those is a causal claim.

THE OUTCOME MEASURE IS CRUDE

In production and used, or not, at eighteen months. It says nothing about whether the deployments that shipped were worth shipping, whether they delivered the value claimed for them, or whether any of them were later switched off.

CODING WAS NOT BLIND

The same people who formed the hypothesis did the coding, on 61 cases, without an independent second coder. This is the weakest part of the method. The coding frame and the case-level codes are available on request specifically so that somebody else can disagree with them.

IT SAYS NOTHING ABOUT UNREGULATED SETTINGS

Every deployment here sits where a wrong output has a named consequence. Consumer products, internal tooling with no external exposure, and anywhere nobody has to accept anything may fail in entirely different places.

The download

Download the full report.

The full report — eight exhibits, the complete method, and every appendix.

Five fields, and the file.

The finding, the method and the limits are on this page. The form is for the complete document.

It arrives in your browser on the next screen, not by email. We ask about your sector because it tells us which sectors read which report, and that is a finding in itself.

Verification widget mounts here.

Questions

Before you read further.

Do I have to give an email to read it?

No — the finding, the method, the summary, all eight exhibits and the full limits section are all on this page. The email address is only for the complete formatted PDF, which adds the appendices and four detailed case accounts.

Is sixty-one enough to conclude anything?

Sixty-one is enough to see a strong, consistent split — 83% production where accountability was designed first, against 18% where it was deferred — but it is not a large or random enough sample to treat as representative of AI deployments generally. Treat it as a documented pattern worth checking against your own experience, not a rate to cite.

Does this apply outside regulated settings?

Not established by this report. Every deployment studied sat where a wrong output has a named consequence — healthcare, banking, professional services, education. Consumer products and internal tooling with no external exposure may fail at entirely different points, and the report's own limits section says so directly.

What counts as 'acceptance' in this report?

The moment an output leaves the system and a named person acts on it without personally re-deriving it — a clinician acting on a classification, an underwriter approving against a score. It's where responsibility for a decision changes hands, and where 26 of the 61 deployments stopped.

Can I get the underlying data?

The coding frame and case-level codes are available on request, specifically so someone else can disagree with them — the report says this itself, since the coding was not independently blinded. Email research@xonik.com.

Who conducted this and when?

XONIK Research, fieldwork January 2024 to June 2026, published August 2026. Thirty-four deployments through direct engagement and twenty-seven through post-hoc review with access to records and staff — not a survey.

Cite this

Govil, A. (2026). Where deployments stop: sixty-one AI deployments in regulated operations. XONIK Research, Report 01.