← All posts

AI in Production

The month after go-live

The deployment worked. Everybody moved on. Four weeks later the decline has already started and nothing has alerted.

PROJECT GO-LIVE OWNED 2 HRS/WEEK PROJECT GO-LIVE Nothing failed. Attention moved.

Nothing failed. Attention moved.

The decline starts on the day the project closes, and nothing marks it.

In short

A project has an owner because it has a budget. When the budget closes the ownership closes with it, and the system does not know that. In the four weeks after go-live, tacit knowledge disperses, the first exception is handled unrecorded, and small drifts start that nothing alerts on — because alerts were built for failure and none of this is failure.

The deployment is live. It works. The team that built it has a retrospective, a handover document, and a new project starting on Monday.

Four weeks later, three things have happened and none of them was planned.

The tacit knowledge disperses. Not the documented parts — those are in the handover. The rest: why that threshold is set where it is, which failure mode is safe to ignore, who to call when the upstream feed is late, what the workaround was for the case that comes up twice a month. That lives in two or three heads, and it does not leave with those people. It leaves with their attention, which is faster and harder to notice.

The first exception arrives and is handled by whoever is nearest. The handling is competent and it is not recorded, so when the same thing happens six weeks later it is handled differently by somebody else. Neither person is wrong. There is simply no accumulating record of what has been decided about this system since the project closed.

Small drifts start. An upstream field changes shape. A volume grows past what was tested. A downstream service updates. Nothing alerts, because alerts were built for failure and none of this is failure — it is a system continuing to work in circumstances slightly different from the ones it was built for.

Alerts were built for failure. What happens after go-live is not failure, which is exactly why nothing catches it.

PROJECT GO-LIVE OWNED 2 HRS/WEEK PROJECT GO-LIVE Nothing failed. Attention moved.
The system is in its best condition on the day everybody stops watching it.

The reason this is invisible is that at go-live everything is working. There is no symptom to point at, no meeting where somebody raises it, and no metric that moves. The system is in the best condition it will ever be in, and that is the moment the attention leaves.

What changes it is small and easy to skip

A named operator with a stated allocation. Two hours a week is typical. Not a team, not a rota — a person whose job includes this system for the next quarter, who knows they have been named.

A thirty-day review, in the diary before launch. Not a check that it is still running. A check of what has been decided about it since, and whether any of those decisions should be written down.

A decision log that outlives the project. So the second exception is handled like the first, by somebody who can see what was done and why.

  • Ownership ends with the budget, and the system does not know that
  • The knowledge that leaves first is the undocumented kind, and it leaves with attention rather than with people
  • Nothing alerts, because none of this is failure
  • Two hours a week, named, is most of the fix

This describes deployments we have been called back into, usually six to twelve months after go-live, which selects for the ones where something had already gone wrong. Deployments that hold may hold for reasons we have not observed.

Run the readiness assessment

Limits

Engagement observation, not a study

Drawn from deployments we were called back into, usually six to twelve months after go-live. That selects for the ones where something had already gone wrong.

Deployments that hold are not represented

We do not see the ones that continue working, so we cannot say what proportion of deployments this describes or what the ones that hold are doing differently.

It describes a particular shape of team

Where the people who built a system are the people who run it, this does not apply. It describes the common arrangement where a project team disbands and an operation continues.

Questions

Is this not what a support contract is for?

Support handles things that break. This describes a system that does not break — it drifts, and drift produces no ticket.

We wrote a handover document. Is that not enough?

Only if somebody has run the system from it. A document nobody has worked from is a document nobody has tested, and the gap between those is where this lives.

How much time does it actually need?

Two hours a week for the first quarter is what we usually see work. The number matters less than that it is stated and belongs to a person.

What if nobody will take it?

That is the useful finding. It means the system entered production without an owner, which is a decision somebody made by not making it.

Does this apply to bought software?

Less sharply, because the vendor handles some of it. It still applies to your configuration, your integrations and your exceptions.

When should this be planned?

Before go-live. Naming an owner after the project has closed means finding budget that no longer exists, which is why it usually does not happen.

Govil, A. (2026). The month after go-live. The Field Report, XONIK.

Written by Amit Govil, Founder, XONIK

More from the Field Report

A fortnightly letter on the distance between deciding and doing.

One piece of research or one working framework, every two weeks.

No sequence, no upsell, unsubscribe in one click.