A screenshot is not a measurement.
Two hundred and twenty-six firms were named at least once. Ninety-four of them appeared in one or two runs out of twenty-six.
155 of 226 appeared in fewer than nine of twenty-six runs.
Being named once is common. Being named consistently is not.
A firm can appear in an AI answer once and never again. Most firms named at all are named incidentally — a single run, a single assistant, a single week. A screenshot proves you were named once. It says nothing about whether you hold a position, and the two are different things.
Somebody in the marketing team asks an assistant to recommend firms in your category. Your firm appears. The screenshot goes into a channel, somebody says the AI work is paying off, and everyone moves on.
The following week the answer is different, and nobody checks.
We ran a fixed prompt set weekly for six months, across four assistants and eleven categories, and recorded every firm named. The distribution is the finding: most firms named at all were named rarely, and a small group was named almost every time. The full method and the numbers are in the report.
What the distribution looks like
Being named once is common. Being named consistently is not, and the gap between them is the whole finding.
The largest group by far appeared in only one or two runs out of the whole period. A much smaller group appeared in nearly all of them. Between those two sits everybody else, and the gap between the ends is the whole point.
Every firm in that group could have produced a screenshot. Most of them, asked the following week, would not have appeared.
Most firms that could show you evidence of appearing in AI answers do not hold a position. They were named once.
Three things one check cannot tell you
Which assistant. Most firms named at all were named by exactly one of the four, and only a handful by all of them. So a screenshot from one assistant most likely describes only that assistant.
Which phrasing. We ran variants a reasonable person would call equivalent. The closest kept most of the baseline’s top five; the most divergent kept fewer than half, meaning the majority of named firms differed on a question that meant the same thing.
Which week. One assistant held its top five fairly steady week to week. Another changed most of its list in a typical week. Both were answering the same question.
What a measurement requires
Four things, none expensive.
- A fixed prompt set, written down. Not one prompt — a set covering how a buyer might actually ask. Fixed, because changing the wording changes the answer, and a baseline you cannot repeat exactly is not a baseline.
- More than one assistant. Three or four. Measuring one assistant and generalising is the commonest error, and given how little the four overlap, a large one.
- Repetition over weeks. A few weekly runs across three or four assistants tell you whether you are in the stable set or the incidental one. A single run tells you about a moment.
- Recording what came back, not just whether you appeared. Which firms were named, in what order, and what was said about them. Your own presence is one data point; the shape of the answer is the useful part.
That is an afternoon to set up and twenty minutes a week to run. Any supplier charging for AI visibility work who cannot show you their prompt set has not produced a measurement, and that is the first thing to ask for.
What separates the consistent from the incidental
The consistently-named firms share one thing the incidentally-named do not.
Almost all of the consistently-named were cited in writing they did not control. Almost none of the incidentally-named were. That gap is wider than any other we measured, and it is the one nobody can buy.
Crawlability was near-universal in both groups. So was schema markup. Publishing a blog monthly was, if anything, slightly more common among the firms nobody names. None of the structural characteristics separated the two groups.
The second-largest gap is whether the firm had published original research. That one separates the groups almost as sharply as citation does.
Why this matters commercially
The single screenshot produces two expensive errors.
Stopping too early. A firm sees itself named, concludes the work is succeeding, and stops. Nothing was established and the position was never stable.
Buying the wrong thing. A supplier runs one prompt, screenshots it, presents it as a baseline, and sells six months of work against it. At the end they run it again. Whether the number moved tells you nothing, because the measurement had no stability in either direction.
We sell this work, which is worth stating plainly. It is also why the finding that most changed our own advice is in Report 03: across the firms doing structural visibility work, six months moved the structural characteristics substantially and moved the outcomes very little.
What to do this week
- Run the check properly once. Four assistants, one category, your baseline prompt. Twenty minutes. Write down every firm named, in order.
- Do it again next week. The difference between the two lists tells you more than either list alone.
- Check your robots.txt. One firm in our set was invisible to every assistant for two years because an agency had added a line in 2024 to block scraping. Nobody at the firm knew. Four minutes, and occasionally the entire answer.
The same problem in a different form — measuring the flattering thing rather than the true one — is covered in what a proof of concept actually proves.
The counts, the method and the limits are in What assistants recommend, free to read in full.
Limits
The models changed under us
At least two of the four assistants shipped significant updates during the study period. We did not control for this, and it is the largest limitation on every stability figure here.
Eleven categories is not a sample
Chosen to span concentration levels, not sampled from any population.
Coding was not blind
Characteristics were coded by hand by the people who expected the corroboration finding. The frame is available on request.
We sell this work
Parts of this run with that commercial interest as well as against it.
Questions
How many runs before I have a baseline?
A few weekly runs across three or four assistants distinguishes a stable position from an incidental one. Fewer, and you cannot tell a real change from ordinary week-to-week movement.
Can I just use a tool that tracks this?
Some are useful. Ask two questions before buying: what is the prompt set, and how many assistants. A tool measuring one assistant on one phrasing has the same problem as the screenshot, at a subscription price.
Does appearing more often bring enquiries?
We did not measure that and this cannot tell you. It measures appearance, not conversion, and anyone claiming a demonstrated link should be asked how they measured it.
Our category has only three names in it. What then?
Concentrated categories are hard to enter and stable once entered. The top few take almost everything, and only a handful of firms are ever named at all. That is a long horizon, worth knowing before you budget for a quarter.
Is this different from SEO?
It overlaps. Some structural work serves both. What differs is that this rests on corroboration and entity resolution, which a traditional audit does not examine and which no amount of on-site work produces.
How long before structural work shows?
Structural characteristics move in months. Outcomes move over quarters, because they depend on other people. Anything promising a quarterly outcome is describing paid media under a different name.
Govil, A. (2026). A screenshot is not a measurement. The Field Report, XONIK.