Try this. Open ChatGPT, ask which company a customer should use for what you do, in your city. Read the answer. Open a new chat and ask the identical question.
The answers will differ. Sometimes mildly — the same names in a different order. Sometimes completely — a competitor who was absent is now first, and a business that was recommended has vanished.
Most people treat this as a flaw and move on. It is not a flaw. It is the single most useful property of the system, because it tells you that AI visibility is not a position you hold. It is a probability, and probabilities can be measured.
Two sources of variation
The first is that the model does not retrieve a stored answer. It generates one, token by token, sampling from a distribution of plausible continuations. Where one candidate name is overwhelmingly more likely than the alternatives, it will appear nearly every time. Where three names are roughly equally plausible, which one surfaces is close to a coin toss.
The second is retrieval. When an assistant runs a live search, it fetches a small number of pages and composes from those. Which pages it fetches depends on how it rewrote your question into a query, what the search index returned at that moment, and how many results it decided to open. Two runs, two slightly different sets of sources, two different answers.
These compound. Variable inputs feeding a sampled output produces answers that can differ substantially without anything having changed in the world.
The number that actually matters
Once you accept the variance, the right measurement becomes obvious: not am I mentioned, but how often.
Those cases are different situations requiring different responses:
Never mentioned, across many runs and all four assistants. You are absent from the material the answers are built from. Nothing incremental will fix this. Something has to change in the sources.
Mentioned occasionally — say, one run in five. You exist in the material but you are not the strongest candidate. This is a far better position than it feels like, because the distance between one-in-five and three-in-five is usually a few sources rather than a rebuild. It is also the position most often misdiagnosed, because whether an owner feels invisible or fine depends entirely on which run they happened to see.
Mentioned most of the time but not always. You are a default answer with a soft edge. Worth knowing which competitor displaces you on the runs where you lose, because that is the specific contest you are in.
Mentioned every time. Strong, and more fragile than it looks. Consistency this high nearly always rests on a handful of sources. It is worth knowing which ones, so you notice when one of them changes.
You cannot land in any of these categories from one check. One check tells you the outcome of one sample, and an owner who asks once and sees their name walks away reassured on evidence that does not support it.
Why one assistant is not four
Layered on top of run-to-run variance is variance between assistants. ChatGPT, Claude, Gemini, and Perplexity differ in training data, in search back-end, in how aggressively they search at all, and in how willing they are to name specific businesses.
The practical result is that being absent from one assistant tells you very little. Being absent from all four is a pattern, and it is a pattern with a single cause: the sources. Being strong in one and absent in three is also informative — it usually means a source that one assistant's search back-end favours is carrying you, and nothing else is.
Phrasing is a third dimension
The same customer intent arrives in many forms. Best dentist in Tijuana, where should I get implants in Tijuana, is it safe to get dental work done in Mexico, cheapest dental clinic near the border. These are the same person at different points in the same decision, and they routinely produce different recommendations, because each rewrites into a different search and matches different material.
A business can be a default answer for one phrasing and invisible for another — frequently invisible for the earlier, more cautious phrasings, which are the ones customers use before they know who the options are. Those are the queries where a recommendation is worth the most, and they are the ones owners are least likely to think to test, because they are not questions about you.
What this means for measuring anything
Put the three dimensions together — repeated runs, four assistants, many phrasings — and the reason casual checking fails becomes clear. You are sampling one point out of a large space and drawing a conclusion about the whole of it.
That is also why the conclusion people reach is so often wrong in a specific direction: reassuring. Ask the assistant about yourself by name, once, on the one you happen to have open, and it will tell you about your business, because you supplied the name. Nothing about that exchange resembles what a customer who has never heard of you experiences.