Run the same prompt through ChatGPT, Claude, or Google's AI a hundred times and your chance of getting the identical list of recommended brands twice is under one in a hundred. Getting that same list in the same order is closer to one in a thousand. That comes from research by Rand Fishkin of SparkToro with Patrick O'Donnell of Gumshoe.ai, published in January 2026, in which 600 volunteers ran 12 prompts across the three tools a combined 2,961 times. The implication is uncomfortable for a fast-growing industry: any product selling you a "ranking position in AI" is reporting a number that does not exist in a stable form. Fishkin's own verdict on such tools was that they are, in his words, full of baloney.
What the study actually did
- Scale: 600 volunteers, 12 prompts, repeated 60 to 100 times per platform, 2,961 responses in total.
- Categories: deliberately varied, from chef's knives and headphones to cancer care hospitals and digital marketing consultants.
- Method: each response normalised into an ordered list of brands, then compared for overlap, order, and repetition.
- Result: the same list almost never appeared twice, the order was almost never consistent, and even the number of items varied between runs.
A second, related finding from the same team is arguably more useful: when 142 people were asked to write their own prompt expressing one shared intent, the average semantic similarity between those prompts was 0.081. People describing the same need phrase it almost nothing alike, which means the "prompt" a tracking tool monitors is one arbitrary phrasing among thousands your customers actually use.
It gets worse when the conversation continues
Real buyers do not stop at one question. Research from Clovion, reported in 2026, ran tens of thousands of multi-turn conversations across Claude, ChatGPT, and Gemini and found that adding a single buyer detail, something as ordinary as "for a small team", removed around 62% of the brands from the first answer.
Worth flagging honestly: Clovion sells AI visibility tracking, so that is a vendor self-study, and the published figures were corrected after release. Treat the direction as plausible and the precision as unverified, the same discipline we set out in our guide to reading search statistics. Even discounted heavily, it tells you something real: the shortlist you appear on in a first answer is not the shortlist a buyer ends up with.
The reassuring half nobody quotes
Here is where most coverage of this research stops, and where it becomes genuinely useful. The same study found that while the exact list is random, the pool of brands is not. Across 994 responses to varied headphone prompts, a handful of major brands appeared in roughly 55% to 77% of answers. In tight categories with few serious players, the same core names kept surfacing in different sequences.
So the picture is not chaos. It is a lottery with weighted tickets. Individual draws are unpredictable; how many tickets you hold is very much not. Your job is not to win a position, it is to be in the pool often enough that you keep getting drawn.
What to measure instead of rank
- Presence rate. Across many runs of similar prompts, in what percentage of answers does your brand appear? That number is stable enough to trend and honest enough to act on.
- Prompt variety, not prompt precision. Track a spread of phrasings that express the same intent, because your customers will not use the one your tool picked.
- Multi-turn survival. Ask the follow-up question a real buyer would ask, then check whether you are still in the answer. Appearing in turn one and vanishing in turn two is a specific, fixable problem.
- Sentiment, not just presence. Being named with caveats attached is a different outcome from being named as the recommendation, and most tracking ignores the distinction entirely.
- Per platform. Results diverge sharply between assistants, so a single-platform check tells you about one system, not about your visibility.
- Mentions versus citations. Being used as a source and being named in the answer are separate achievements, as covered in our piece on ghost citations.
An honest correction to our own advice
Our 15-minute AI visibility self-test asks you to run each question once, in a fresh session. Given this research, once is not enough for the recommendation questions specifically. Run those three times each and record how many times you appear rather than whether you appeared, then compare that fraction month to month.
It adds about five minutes and it converts a coin flip into a measurement. We would rather update the method publicly than leave you measuring noise.
Try it yourself in two minutes
Open a fresh session with any assistant, ask it to recommend the best businesses in your category and city, and note the names. Open another fresh session, ask the identical question, and compare. Then add one detail a real customer would add and watch the list rearrange again. Most people who run this small experiment stop trusting single checks immediately, which is the entire point.
Frequently asked questions
Does this mean AI visibility tools are useless?
It means rank-position metrics are unreliable, not that measurement is impossible. Tools that report presence rate across many runs and many phrasings are measuring something real. Tools that report "you are ranked third in ChatGPT" are reporting a single draw from a random process.
Why do the answers change at all if my content has not?
These systems generate probabilistically rather than looking up a stored ranking, so variation is built into how they produce text. The underlying associations are more stable than any individual output, which is exactly why presence rate works better than position, and why the mechanics in our guide to where AI answers come from matter more than any single result.
How many runs do I need for a meaningful number?
The research suggests patterns emerge somewhere in the region of 60 to 100 runs, with narrow categories stabilising faster than broad ones. For a small business doing this manually, three to five runs per question monthly is enough to spot a trend, provided you compare fractions rather than single outcomes.
So what actually increases my presence rate?
The same unglamorous work as ever: consistent identity across the web, original material only you can provide, independent mentions by name, and clean structure. Those raise how many tickets you hold in every draw, which is the whole argument of our AEO playbook.
The bottom line
AI recommendations are a weighted lottery, not a ranking. Stop asking where you rank, start asking how often you appear, run the check several times before believing it, and spend your effort on the things that put more tickets in the drum. If you want a baseline measured properly across multiple runs, phrasings, and assistants rather than a single lucky screenshot, our free SEO and AEO audit does exactly that, and reports the fraction rather than the flattering example.