Research

What assistants recommend

A fixed prompt set run weekly for six months across four assistants, in eleven professional-services categories.

Two hundred and twenty-six firms were named. Thirty-three appeared consistently, and their websites are not why.

Run the visibility assessment

The finding

Across 264 prompt runs in eleven categories, 226 distinct firms were named at least once and 33 appeared in seventeen or more of the twenty-six weekly runs. The separating characteristic was not site quality, technical SEO or content volume — those were near-universal. It was third-party corroboration: 86% of consistently-named firms were cited in writing they did not control, against 9% of firms named once or twice.

LengthEight exhibits, full method and appendices
Basis264 runs, 4 assistants, 26 weeks
AccessFree, in full

What is in it

The argument, in full.

Not a preview. What the study looked at, what it found, and what separated the outcomes.

A fixed set of prompts was run weekly, at the same time on the same day, against four consumer AI assistants, across eleven professional-services categories, for twenty-six consecutive weeks from February to August 2026 — 264 runs in total. No personalisation, no memory, no prior context carried between sessions; each run started clean. The study exists because most published advice on appearing in AI-generated answers is transposed from search engine optimisation on the assumption the mechanisms are similar, without anyone having actually watched what these answers contain over a sustained period. XONIK has a commercial interest in the answer, which the report states plainly rather than leaving implicit.

Two hundred and twenty-six distinct firms were named at least once across the study. That number alone overstates how meaningful "appearing in AI answers" is: 94 of those firms were named in only one or two of the twenty-six runs, and being named once produced no guarantee of being named the following week. Only 33 firms — about one in seven — appeared in seventeen or more of the twenty-six runs, and just 7 were named by all four assistants on the same question in the same week. The four assistants disagreed with each other more than they agreed: 167 of the 226 named firms were named by exactly one assistant.

The report tested six characteristics against the gap between consistently-named and incidentally-named firms, and most of what's commonly recommended turned out not to be the answer. Technical crawlability, directory listings and schema markup were near-universal in both groups, meaning they're a threshold to clear rather than a lever that moves anything once cleared. What separated the two groups was third-party corroboration — being cited in writing the firm didn't control, such as trade press or industry publications — present in 86% of consistently-named firms against 9% of incidentally-named ones. Publishing original research showed the second-largest gap, 71% against 6%, and is plausibly the mechanism behind the citation gap rather than a separate lever, since firms that publish something worth citing tend to get cited.

Thirty-one firms in the wider set underwent deliberate structural visibility work — entity consistency, passage-level restructuring, directory listings — tracked before and after six months. The inputs moved substantially: entity consistency rose from 52% to 89%, extractable passages from 44% to 81%. The outcomes barely moved: citation rose from 9% to 14%, and appearing in nine or more runs rose from 12% to 17%. The report states this finding runs against its own commercial interest in recommending structural work, offered as a reason to trust it, while flagging that the sections built partly on the same underlying data are a reason to check them independently.

How much the question itself matters was also tested: four prompt variants run alongside the baseline retained between 39% and 62% overlap with the baseline's top five firms, meaning a visibility measurement built on one phrasing is a measurement of that phrasing, not of visibility generally. The report's own conclusion is that structural work is necessary and achievable but not sufficient — the ceiling sits with corroboration, which is produced by other people writing about a firm, not by anything a firm can do to its own site alone.

The exhibits

Eight charts, all in the PDF.

01

Distribution of firms by how many runs named them

02

Characteristics present, consistently-named against incidentally-named firms

03

Named firms by how many of the four assistants named them

04

Week-to-week stability of each assistant's top five

05

Share of mentions taken by the top three firms, by category

06

Rank position over twenty-six weeks, one category, one assistant

07

Overlap with the baseline prompt, by variant phrasing

08

Characteristics before and after six months of structural work

Method

How the finding was reached.

Period February to August 2026, 26 consecutive weeks
Assistants Four consumer assistants, anonymised as A to D. Named in the working material.
Categories 11 professional-services categories, UK and US framings run separately
Runs 264 — 4 assistants x 11 categories x 6 monthly sampling points, plus weekly runs on 3 categories
Prompt One baseline prompt per category, plus four variants tested in weeks 9 and 18
Session Fresh session per run. No memory, no personalisation, no location signal beyond the default.
Recorded Every firm named, in order, with the surrounding sentence
Selection Not representative. Categories chosen to span concentration, not sampled from a population of categories.

Limits

What this does not show.

Written before the analysis was run. Afterwards, limits come out shaped to protect what was found.

THE MODELS CHANGED UNDER US

At least two of the four assistants shipped significant updates during the twenty-six weeks. We did not control for this and could not. It is the largest limitation on every stability and drift figure in the report, and it means the volatility we measured is a mix of genuine week-to-week variation and step changes we cannot separate.

ELEVEN CATEGORIES IS NOT A SAMPLE

The categories were chosen to span concentration levels, not sampled from any population. The category-level findings describe these eleven and should not be generalised to a twelfth without checking.

CHARACTERISTICS WERE CODED BY HAND, UNBLINDED

The same people who expected the corroboration finding coded the characteristics that produced it. On "cited in third-party writing" in particular, a coder who expects a result can find it. The coding frame and the firm-level codes are available on request and a re-code would be genuinely useful.

WE SELL THIS WORK

We have a commercial interest in the conclusion that structural visibility work matters. The finding in section 08 — that six months of it moved outcomes by five points — runs against that interest, which is one reason to trust it. Sections 03 and 09 run partly with it, which is a reason to check them.

CORRELATION, THROUGHOUT

Nothing here establishes that citation causes appearance. Firms that get cited may share other properties we did not measure. Firm C in section 11 is the closest thing to a natural experiment in the set, and it is one firm.

The download

Download the full report.

The full report — eight exhibits, the complete method, and every appendix.

Five fields, and the file.

The finding, the method and the limits are on this page. The form is for the complete document.

It arrives in your browser on the next screen, not by email. We ask about your sector because it tells us which sectors read which report, and that is a finding in itself.

Verification widget mounts here.

Questions

Before you read further.

Do I have to give an email to read it?

No — the finding, the method, the summary, all eight exhibits and the full limits section are all on this page. The email address is only for the complete formatted PDF, which adds the full prompt instrument and four detailed firm accounts.

Is 264 prompt runs across four assistants enough to conclude anything?

It's enough to see a real, sizeable gap — 86% of consistently-named firms cited in writing they don't control, against 9% of firms named once or twice — but the eleven categories were chosen to span concentration levels, not sampled from the wider economy, so a twelfth category shouldn't be assumed to behave the same way.

Does this apply to any AI assistant, or just the four tested?

Only the four tested, run under one fixed prompt set. At least two of the four shipped significant model updates during the study, which the report treats as its largest limitation — what any other assistant does with a different prompt isn't something this data can speak to.

What actually separated the firms that got named consistently?

Not technical SEO, site quality or content volume — those were near-universal in both groups. The gap was third-party corroboration: being cited in writing the firm didn't control, such as trade press, industry bodies or other firms' publications.

Can I get the underlying data?

The full prompt set, the assistant identities, all 264 recorded responses, and the firm-level characteristic codes with firm names removed are available on request. Email research@xonik.com.

Who conducted this and when?

XONIK Research, run weekly from February to August 2026 across 26 consecutive weeks, published August 2026. Every run used a fresh session with no memory or personalisation carried over.

Cite this

Govil, A. (2026). What assistants recommend: 264 prompt runs across four assistants and eleven categories. XONIK Research, Report 03.