Plugin Visibility Index · Run 04

Asking the same question twice measures the cache, not the engine

Put an identical query to Perplexity twice in one day and it returns the identical source list — every domain, both times. Add a question mark to the end and a median of 21% of the sources change. That single character is the difference between measuring a cache and measuring an engine.

Identical query
100%
One character added
79%
Different engine
12%
Engine
Perplexity

Why this run exists

Run 03 found that Perplexity and Google AI Mode share a median of only 15% of their sources for the same question. It could not say how much of that was a real difference between the engines and how much was ordinary churn — the same engine simply returning different pages each time you ask.

The obvious way to find out is to ask one engine the same question twice and see how much its own source list moves. That is what I did, and the result was not what I expected.

What happens when you repeat the exact query?

Nothing. In both categories I tested, the second run returned the same source list as the first — ten of ten domains for backup, twelve of twelve for membership. Not similar. Identical.

The two runs were separate searches with different thread IDs, hours apart. The answer was regenerated; the retrieval was not. That is a cache keyed on the query string, and it means a monitoring routine that submits the same prompt every week is measuring its own repetition.

If you are buying AI visibility monitoring

Ask the vendor whether their prompts vary between runs. If the wording is fixed — which is the natural way to build a tracker, because it looks like a controlled measurement — a flat, stable chart may be showing you a cached retrieval rather than your actual position. Reassuring and empty.

The full buyer’s version of this — which tracked metrics the caching degrades, which survive it, and the five questions to put to a vendor — is on its own page: Are AI visibility tracking tools accurate?

How much does one character change?

I added a question mark to the end of the same five queries. Nothing about the question changed for a reader. The retrieval changed in every one.

CategorySources beforeSources afterOverlap
Translation141487%
Popups101082%
Membership121379%
Caching101067%
Backup101742%

Median 79%. So roughly a fifth of an answer’s sources are not fixed by the question — they are decided by the exact string you typed.

So is the cross-engine difference real?

Yes, and by a wide margin. Stacking the three comparisons on the same five categories:

What variesMedian source overlapWhat it measures
Nothing — identical string100%The cache
One character of wording79%Retrieval churn within one engine
The engine12%A genuinely different index

Rewording a question moves about a fifth of the sources. Changing engine moves about seven-eighths of them. Whatever is different between Perplexity and Google AI Mode is roughly four times larger than the noise inside either one, which is the answer Run 03 was missing.

The practical version: if you are cited on one engine and not the other, that is a real gap and not bad luck. It is worth fixing separately.

What happened to the backup answer?

Backup was the outlier at 42%, and the reason is worth its own section. Adding the question mark took the source count from ten to seventeen, and six of the seven new domains were locale mirrors of wordpress.org:

Host citedLocale
wordpress.orgcanonical (English)
de.wordpress.orgGerman
ko.wordpress.orgKorean
pl.wordpress.orgPolish
tw.wordpress.orgChinese (Taiwan)
hr.wordpress.orgCroatian
br.wordpress.orgBreton

Seven wordpress.org hosts in one answer to an English question, six of them translations of the same plugin directory. Attribution 01 first noticed this behaviour and Run 03 saw it again in two other categories. This is the largest instance so far, and it appeared because of a question mark.

For a site with hreflang alternates, that is the risk in concrete form: one phrasing consolidates on your canonical, another spreads across six translations of it. You do not control which phrasing a buyer types.

What can this run not tell you?

Limits, stated plainly

Small samples, and I would rather say so than round them up. The identical-query test is two categories. The question-mark test is five. One engine, one day.

Two categories returning identical lists is strong evidence of caching but does not tell you how long the cache lives, whether it is per-account, or whether it applies to all query types. A weekly monitoring routine might fall outside it. I have not measured that, and anyone relying on this should test their own interval rather than take 100% as a general law.

The question mark is one specific edit. It is a good test because it changes the string without changing the meaning, but it is not a general measure of how much rewording moves retrieval — a different word, a plural, a year removed, might move more or less.

The 12% cross-engine figure here is the five categories tested in this run, not the ten in Run 03, which is why it differs slightly from the 15% published there. Both are the same measurement on different subsets.

What happened next

Run 05 adds ChatGPT as a third engine, and finds it displays seventeen per cent of what it retrieves.


Is your tracking measuring anything?

Send me your wordpress.org slug and the prompts you or your tool are tracking. I will run them with and without variation and show you which numbers move and which are the cache. Free, no call, usually within a day.