The visibility index · 2026 edition
Ask it again.
Your buyers have started asking a machine instead of searching. The machine answers with total confidence, names a short list of companies, and moves on. You are not in the room for any of it.
An entire category of tools has appeared to tell you how you fare in that room. Almost all of them ask the machine once.
One
Nobody can see the prompts.
Not us. Not anyone. There is no API that returns what people typed into an assistant, no log, no panel. The companies operating these models do not publish it and have no reason to.
So when a tool offers to show you the prompts people ask about your brand, it is modelling from search data and presenting the output as observation. That is not a small liberty. It is the difference between a measurement and a guess wearing a chart.
We say this first because a company selling measurement cannot overstate its own instruments and expect to be believed about anything else.
What can be observed is the answer itself. So we observe the answer itself.
Two
We asked the same questions 7 times each.
16 questions across 4 markets, from wide open to tightly specific, each put to an assistant with live web search 7 separate times. Every company named in every answer was recorded. Nothing was cleaned up.
Averaged across the 16 questions
Ask once and you see fewer than half the companies the same model names across 7 asks. Only one in 6 companies turns up every time. Nearly half turn up once and never again. Nothing changed between those asks. Not the question, not the market, not the model. Only the roll.
| Question shape | Seen in one ask | Named every time | Same company first | Top three repeated |
|---|---|---|---|---|
| Open category | 40% | 11% | 71% | 53% |
| Ranked recommendation | 51% | 24% | 75% | 67% |
| Specific use case | 41% | 14% | 93% | 57% |
| Single attribute | 40% | 14% | 75% | 42% |
Narrower questions produce shorter lists, as you would expect. They are barely more consistent. Even on the steadiest shape, only a quarter of the companies named turned up in all 7 asks. Tightening the question shrinks the list without settling it.
The finding that argues against us
The leader is fairly durable. Across every question, the company named first held that position in 79% of asks. If all you want to know is who is winning, one ask will usually tell you.
The podium is not durable. Ask twice and two top-three lists share barely half their members (55%). We publish this because a paper that only reports the numbers supporting its author is an advertisement.
Three
A number without a distribution is not a measurement.
Here is the consequence, and it is larger than it first sounds. If a tool asks once, it has a single point and no idea how far that point moves on its own. When the figure changes next month it cannot tell you whether you gained ground or whether the dice landed differently.
Every claim of improvement made on a single sample is unfalsifiable. Not wrong necessarily. Unfalsifiable, which is worse, because it cannot be checked by anyone including the person making it.
You cannot know that something moved until you know what standing still looks like.
That is the entire reason to ask repeatedly. Not thoroughness. Not diligence. You are establishing the width of the noise so that a real change can be told apart from it.
Four
Being mentioned is a symptom.
Underneath the mention sits something more durable: what the model holds to be true about you. What you sell, who you serve, what you cost, who you sit beside. The mention is downstream of that belief.
This matters because you cannot make a model mention you. There is no lever. You can correct what it believes, and the mentions follow. A category that measures only presence is watching the symptom and prescribing for it.
It also reframes what is at stake. Being absent from an answer is a marketing problem. Being described wrongly in an answer, to a buyer, at the moment of decision, is not a marketing problem.
Five
Your own site is often the source of the error.
A post published three years ago explaining that you serve one kind of customer will keep saying so long after it stopped being true. It may be your best performing page. Every tool that measures content by traffic will report it healthy while it quietly teaches machines something false about you.
There is a second, less obvious problem. Deleting it does not undo it. A page removed today can persist inside a model for a year or more, because retrieval and training are separate systems with separate clocks. Pruning the page removes your ability to correct the record while leaving the belief in place.
The remedy is not deletion. It is publishing something clearer, better sourced and easier to retrieve than the thing it replaces.
Six
There are two rooms now, not one.
In the first, a person asks and reads an answer. Success is being named, placed well and described accurately. That is the room everyone is currently measuring.
In the second, software acts on someone's behalf. It does not read your paragraph. It needs to locate you, parse what you offer, compare it against alternatives on terms it can evaluate, and complete something. Prose is not an interface.
Being recommended and being purchasable are different achievements.
A company can be famous in the first room and unreachable in the second. That is not a ranking problem and it will not be solved by writing more. It is answered by publishing what you sell in a form a machine can act on, and it is testable today: give an agent the task and see whether it completes.
Seven
What we hold ourselves to.
Every question carries its source. A question observed in your own search data and a question a model invented are both useful and they are not the same kind of thing. They are labelled differently and never blended into one score. Blending provenance is how a guess becomes a number.
Observed and modelled stay visibly apart. What an assistant actually said is captured word for word and shown to you. What we infer is marked as inference.
No figure appears that does not trace to something real. Including in this paper. If a measurement is too thin to support a claim, it reads as unavailable rather than as a confident number.
We will tell you when we cannot prove your work paid off. This is the least commercially convenient sentence in the document and the one we are least willing to remove.
How the measurement was made
16 questions were generated across 4 markets, in four shapes ranging from open to specific. Each was asked 7 times through a single frontier model with live web search enabled, using the same pipeline that runs a customer audit. Company names were extracted from each answer by a separate, deliberately minimal pass with no knowledge of any target company, so the extractor could not be primed to find one. Names were normalised for case and legal suffix only.
Limits worth stating. One model and one provider, so this measures that model rather than the category. An API is not the consumer product it shares a name with, so what a buyer sees may differ. All asks fell within a single day, which measures variance between asks rather than drift over time. Absent measurements from failed calls were excluded rather than counted as absences.
Ask once and you get an answer. Ask again and you get the truth.