Methodology
How TendorAI measures AI visibility, stated so that someone else could run the same measurement and check us — including the parts this design cannot reach.
What is measured
Whether AI assistants name a business in their answers to a defined set of questions, how they describe it when they do, and which sources they cite.
That is the entire object of measurement. Not how well a website is optimised, not a model of how assistants are thought to behave, and not an estimate derived from public signals. A question is put to an assistant, an answer comes back, and the answer is recorded.
Everything else in this document is about doing that carefully enough for the result to mean something.
Question and prompt design
A panel is a fixed list of prompts, written as a buyer would actually phrase them, fixed before collection starts and not altered afterwards.
Our published solicitor studies use a panel of 68 prompts across 17 UK cities, with four prompt types per city:
- Best-in-city — "Best conveyancing solicitors in Bolton"
- Purchase intent — a prompt phrased as someone ready to instruct
- Reputation — a prompt asking about standing rather than availability
- Practice-area specialism — a prompt naming a specific area of work
The specialism prompts span seven practice areas. Those seven apply to the 17 specialism prompts, not to the whole panel — a distinction we originally got wrong and corrected publicly on 02/08/2026, with the panel size, run count and citation figures unchanged. The full panel is published as a CSV, so the wording of every prompt can be read rather than taken on trust.
For client measurement the same principle applies: the question set is agreed at onboarding, fixed, and kept constant between measurements. Changing the wording later would move the numbers for reasons that have nothing to do with the business, which would make the comparison worthless.
Repeated runs
AI assistants do not return the same answer to the same question every time. A single run is therefore a sample of one, and reporting it as a position would be a basic error.
So each prompt is run several times and the results are aggregated. In the August 2026 study each of the 68 prompts was run ten times. Each of 1,214 firms was assessed against the four prompts for its own city at ten repeats — 40 eligible answers per firm, and 48,560 firm-answer observations in total.
Repetition is what converts “the assistant mentioned you once” into a statement about consistency, which is usually the more useful fact.
Which AI assistants
Every measurement names the assistants it used, because they do not behave alike and a result from one is not a result from another.
- Client measurement covers ChatGPT, Google AI Overviews and Perplexity.
- The July 2026 study used two engines, Perplexity and ChatGPT.
- The August 2026 study used one, Perplexity.
The July study says plainly in its own limitations that Google AI Overviews reaches more UK users than either engine it covered and is not represented in it, and that its findings should not be generalised to AI search as a whole. That is the correct way to read any of this: a result belongs to the assistants it was collected on.
Recorded responses and citation capture
Every answer is recorded and kept. Figures are calculated from those records, so any number can be traced back to the responses it came from.
Where an assistant shows the sources behind an answer, those are captured too. In the July 2026 study 12,279 citations were recorded and each was classified by domain type against a published classification list of 52 domains, so the classification can be inspected and disagreed with.
Citations are recorded as sources cited in the response. They are never described as sources that influenced it. A citation establishes that a source appeared; it does not establish that the answer was drawn from it, and this design has no way of testing that.
The units a result is expressed in
Precision about units matters, because the same underlying data can be reported in ways that sound very different.
- Firm-answer observation — one firm, assessed against one answer in which it was eligible to appear. The base unit.
- Eligible answers — the answers a given business could have appeared in, given the prompts covering its city or market.
- Mention rate — the share of eligible answers in which a business was named.
- Named at least once — whether a business appeared in any of its eligible answers. A threshold, not a ranking.
- Citation — one source appearing in one recorded answer.
- Collection window — the dates between which the answers were collected.
Movements in a rate are reported in percentage points, not as a percentage change, and the underlying counts are given alongside. A rise from 3.59% to 4.19% is 0.60 percentage points; describing the same movement as “up 17%” would be technically arguable and practically misleading.
Baselines and re-measurement
The first measurement of a fixed question set is the baseline. Every later measurement of that same set is compared against it, and the comparison holds only because the set did not change.
Where a comparison is used to assess whether work made a difference, a control is what makes it meaningful. In the August 2026 study we measured a control group of firms we changed nothing about between waves: their mention rate moved from 3.59% to 4.19% over five weeks, up 0.60 percentage points, for no reason we caused.
That figure is the most useful thing in either study. Visibility drifts on its own, so any movement has to be read against that drift before it can be called an effect. Work that beats it may have done something. Work that does not has not been shown to.
How results are interpreted
Observation and interpretation are reported separately, and in that order. What the recorded answers contained is stated first, as fact about a sample in a window. What we judge it to mean is labelled as interpretation, so a reader can accept the first and reject the second.
Hypotheses are labelled as hypotheses and are not used as the basis of a recommendation until they have been tested. Unknowns are stated rather than left for the reader to notice.
Results are also kept inside their population. A finding about SRA-regulated solicitors in 17 cities is a finding about SRA-regulated solicitors in 17 cities. It is not evidence about accountants, about businesses generally, or about how “AI” behaves.
Limitations of the design
- Assistant coverage. Each study covers only the assistants named in it, and they do not behave alike.
- Panel coverage. A panel is a sample of the questions people ask, not all of them. Results describe the panel.
- Population. Our published research covers SRA-regulated solicitor firms in 17 UK cities. It is not a sample of UK businesses.
- Time. Assistants change without notice. Every result describes its collection window and may not hold outside it.
- Variability. Repeats reduce the effect of run-to-run variation; they do not remove it.
- Observation only. We see outputs, not the systems producing them.
- Single target. Client measurement currently measures one business — the client. Competitor measurement is not part of what the instrument does today.
What this methodology cannot establish
Worth stating separately from the limitations, because these are not gaps to be closed by collecting more data. They are outside what this design can do at all.
- That anything caused an answer. The design records what assistants said; it does not isolate why they said it, and it has no mechanism for doing so.
- That a cited source influenced the response it appeared in. A citation is evidence of appearance, nothing more.
- That a change made to a business produced a change in a later measurement. The two can be reported side by side; the causal link between them is not measured.
- How assistants work internally. We observe outputs. We have no access to what happens between the question and the answer.
- Anything about questions outside the panel, markets outside it, or assistants not used in that run.
- Anything about a period other than the stated collection window.
The July 2026 study puts the same point in one line: it is a measured snapshot, not a causal study. It tells you who AI cites, not what makes AI cite anyone.
Changes, deviations and corrections
Every departure from a pre-registered plan is recorded in that study’s deviations log and published alongside it, including corrections to our own errors. The 02/08/2026 correction described above is in the July log.
Methodology is not changed to improve a result or to fit a commercial position. Where a method does change, the change is dated, the reason is given, and earlier figures are restated alongside the new ones so that editions stay comparable.
Published research URLs are permanent. Corrections and new editions are published without breaking existing citations.
The studies, their deviation logs and the datasets behind them are at Data & Evidence. The standards our public claims are held to are on Evidence & Standards.
Read the studies this method produced
Published in full and ungated, with the prompt panel, the classification list and the deviation logs available for anyone who wants to check the work.