Guide 10 · Evaluation
Are AI visibility tracking tools accurate?
They measure something real. Whether it is the thing you think you are buying turns on one design choice that nobody advertises: whether the prompt wording changes between runs.
Fixed wording is the published best practice across this category. And a fixed query string, put to Perplexity twice, returned an identical source list both times — ten of ten domains in one category, twelve of twelve in another, hours apart, on separate searches.
What do these tools actually do?
The shape is consistent across the category. You define a set of prompts your buyers plausibly ask. The tool submits them to several engines on a schedule and records whether your brand is named, whether your domain is cited, where you appear in the answer, and how you compare with competitors. You get a chart over time.
The prompts are meant to stay fixed. That is not a criticism of any one vendor — it is the stated methodology. One vendor’s own guidance puts it plainly: track “a fixed set of buyer prompts” across engines over time.
And the reasoning is sound. If you want to know whether a number moved, you hold everything else still. Changing the wording between runs would mean you could not tell a real change from a rewording. Fixing the prompt is what makes it a measurement rather than an anecdote.
So what is the problem?
Holding the string still can hold the answer still too.
I put five identical queries to Perplexity in separate searches with different thread IDs, hours apart, then added a single question mark to each and ran them again. Three levels of variation, one engine:
| What varies | Median source overlap | What you are measuring |
|---|---|---|
| Nothing — identical string | 100% | A cache |
| One character of wording | 79% | Retrieval churn inside one engine |
| The engine | 12% | A genuinely different index |
The answers were regenerated. The retrieval was not. That is behaviour consistent with a cache keyed on the query string — and a monitoring routine that submits the same prompt every week is, on those queries, re-reading its own last result.
A flat line then means one of two very different things: your position is stable, or nothing was measured. Those look identical on a chart, and the reassuring one is the one you will assume.
Which of the numbers does this affect?
Not all of them, and this is where I want to be careful rather than sweeping. Caching retrieval does not freeze everything a tracker reports — the answer text is still generated fresh each time, so metrics drawn from the prose can still move while the source list sits still.
| Metric | Comes from | Affected? |
|---|---|---|
| Domains cited / citation share | The retrieved source list | Directly |
| Share of voice by source | The retrieved source list | Directly |
| Brand named in the answer | Generated text, drawn from those sources | Bounded by them |
| Position within the answer | Generated text | Can still vary |
| Sentiment / tone | Generated text | Can still vary |
So the citation-side metrics are the ones to distrust on a fixed prompt, and they are usually the ones being sold as the headline. If a page cannot enter the source set, it cannot be cited — and if the source set is frozen, a genuine improvement on your site has no route into the number.
You publish, you improve, you wait. The chart does not move, so you conclude the work is not landing and you stop. The work may have been landing perfectly well into a query the tracker was not really re-asking. A measurement that cannot detect success is worse than no measurement, because it produces a confident wrong decision.
What should you ask a vendor?
Five questions. None of them is hostile and all of them have short answers, which is itself the test — a vendor who has thought about this will answer in a sentence.
| # | Ask | Why |
|---|---|---|
| 1 | Does the exact prompt string change between runs? | The whole question. If no, ask what they have done to verify the retrieval is fresh. |
| 2 | Do you record the source list, or only whether my brand was named? | Different metrics, different reliability, usually one price. |
| 3 | Is each engine reported separately? | Two engines share a median 15% of their sources. A blended score hides the gap that matters. |
| 4 | Have you tested your own tool for cached retrieval? | Submitting one query twice and diffing the sources takes ten minutes. |
| 5 | What would a false stable reading look like in your dashboard? | A vendor who can describe their own failure mode has looked for it. |
You can also run the check yourself in about ten minutes, without any tool. Put one of your tracked prompts to Perplexity, open the full source list and write down the domains. Do it again a few hours later in a fresh search. If the list is identical, you have learned something about your own tracking that no vendor page will tell you.
What are the free checkers measuring?
A different thing, and it is worth separating them out. Search for a free AI visibility checker and you will find around nine on the first page. They split into two kinds.
The first kind takes a brand name and reports mentions across engines. That requires a large stored corpus of prompts and answers — one vendor states 460M+ monthly prompts behind theirs. Whatever else is true, that is a real dataset and the number means something.
The second kind takes a URL and scores your access setup — typically robots.txt rules for a list of AI crawlers, whether an llms.txt file exists, and sometimes a live probe per user-agent. One of the more careful ones names its three signals as exactly that: robots.txt, llms.txt, and a live probe.
Two of those three measure whether an engine can reach you, which is the right question. The third measures whether you have published a file that, on the best available evidence, nothing requests — Ahrefs found 97% of llms.txt files across 137,000 sites had never been fetched by any crawler. The detail is here.
If a free report has told you your AI visibility is fine, check which of the things it scored are things an engine actually asks your server for. And note what none of them check: whether your certificate is valid, whether your server returns a 200 to the retrieval crawler after the firewall has had its say, and whether your content is in the HTML that arrives.
Those three are what I built the free retrievability check to test, which is the fair disclosure here: I have a tool in this category and it is the reason I went looking at the others.
What I am not saying
That these tools are a scam, or that you should not buy one. Tracking is the right instinct — you cannot manage what you do not measure, and the category exists because the need is real.
The claim is narrower: a fixed-prompt design has a specific, checkable failure mode on at least one major engine, most vendors have not published a test for it, and the metrics it degrades are the ones marketed hardest. That is a question to put to a salesperson, not a reason to skip the category.
What this cannot tell you
The identical-query test is two categories. The question-mark test is five. One engine, one day. Two categories returning identical lists is strong evidence of caching but says nothing about how long the cache lives, whether it is per-account, or whether it applies to every query type.
A weekly interval might fall outside it entirely. I have not measured that, and nobody should take 100% as a general law — test your own interval rather than assuming mine transfers. That test is ten minutes and it is specific to you, which makes it worth more than my number.
This is also Perplexity only. Google AI Mode and ChatGPT may behave differently, and I have not run the equivalent test on either. Extending it is the next measurement rather than something I will assert now.
And the question mark is one specific edit. It is a good test because it changes the string without changing the meaning, but it is not a general measure of how much rewording moves retrieval.
Is your tracking measuring anything?
Send me your wordpress.org slug and the prompts you or your tool are tracking. I will run them with and without variation, on both engines, and show you which numbers move and which are the cache. Free, no call, usually within a day.