Somebody has shown you a number. They typed a handful of questions into an assistant, screenshotted the answers, counted the times your brand appeared, and called it your AI visibility. It looks reassuring, and you have no way to tell whether it means anything.
That instinct is right, and the reason is more specific than the usual complaint about AI being unpredictable.
What the number is a property of
Traditional rank tracking starts from a fixed query and observes position against that query. AI-answer visibility is more context-dependent: the answer can change with the wording of the question, the details the buyer supplies, and the system producing the answer.
So the useful question is not what rank you hold. It is: in which buyer situations do you appear, under a defined measurement protocol, and in which do you not?
The score is a property of the measurement protocol, not an intrinsic property of your brand. The questions are the heart of that protocol, and it also needs the rest: what counts as an appearance, which surface or model gets checked, the language and market being tested, and repeated runs under comparable conditions. Change any of those and the same brand produces a different number, honestly.
Decide what you are actually counting
Before running anything, decide what observable event you are measuring. A brand can be mentioned by name. Its page can be cited as a source. It can appear as one option among several. It can be recommended directly.
Those are four different outcomes, and combining them into a single visibility score produces a number that means four things at once.
For a simple primary measure, use brand-name presence: does the brand itself appear in the answer for this buyer situation? Record source citations, comparisons and recommendations separately. A recommendation is worth knowing about, and it is worth knowing about as a recommendation, not as a point inside an aggregate.
The Buyer Question Set
This is the instrument: a small, fixed set of questions written as buyer situations, not as topics. You keep it and re-run it, rather than sweeping broadly once.
Each entry carries three things, and optionally a fourth:
- Buyer situation — who is asking and what constraint they are under
- Exact question — the wording that gets used, unchanged between runs
- Observed result — whether the defined event occurred on that run, plus any stronger outcome you are recording separately
- Surface and language — which assistant, in which language, where that varies
The situation field is what stops the set drifting back into topics. A topic is project management software. A situation is a contracting firm of a dozen people that needs Arabic invoices and has already rejected two tools on price.
What the broad question can hide
What follows is a clearly labelled hypothetical, written for this article. No real brand, and no measured result.
A bilingual accounting practice checks how it shows up. The broad question — best accounting firms for small business — returns a list, and the practice is on it. The number goes in the report.
Then somebody asks the question a real prospect would ask: which accounting firm can handle a locally registered company with a foreign parent and reporting in both Arabic and English? The answer names three practices, describes what each one handles, and leaves them out.
Both questions concern accounting firms. Only one carries the buyer’s actual constraints.
A question set built only from generic prompts can overstate useful coverage, because it removes the specificity that often changes the answer. Size, sector, constraint, language, location — the principle is to include the specificity that materially changes the buying situation, which is rarely all of them, and rarely none.
One run tells you very little
Answers can vary between runs, accounts and system conditions. That makes a single observation a weak basis for a durable conclusion.
Repeating the same small set under comparable conditions lets you see whether the same result recurs. A broad one-off sweep answers a different question: it gives you more prompts once, where a maintained set lets you compare the same buyer situations over time.
Repeated observations can reveal a pattern in what you checked. They do not turn that pattern into a probability of future inclusion.
What a gap actually tells you
When the set shows you absent from a situation that matters, that is information about the answer, and it is worth separating from what you might do next.
A gap has a shape. You may be absent entirely from the situation. You may be present but described in terms that do not match your positioning. You may appear alongside alternatives without being the option the answer recommends. Those are different observations and should remain separate in the record.
No amount of optimisation can guarantee future inclusion in an answer produced by a system you do not control. A visibility measurement should therefore report what was observed, not promise what the next answer will contain.
Start with a small set of buyer situations you actually sell into. Write the question each buyer would plausibly ask, define what counts as presence, and repeat the same questions under comparable conditions at separate times.
What you get is a bounded picture of where your brand appeared in the situations you chose to test — not a universal visibility score. That is far more useful than a number whose question set and measurement rules are hidden.
Before you bolt on another tool, it is worth knowing whether your business runs on systems or on you. I put together a free 2-minute assessment that gives you a straight read on exactly that, and the first thing to fix. Take the free assessment.
Ready to make your AI actually reliable?
Book a diagnosis and we will map the highest-leverage fixes for your business.
Book a diagnosisSharper signal. Smarter decisions.
Join our newsletter for our best thinking on AI and systems, delivered straight to your inbox - no noise.


