Ask the same question twice, get different sources. The volatility data
Ask an AI engine the same question twice and you will often get two different sets of sources. That is not a bug you are catching, and it is not your competitors outmanoeuvring you overnight. It is the defining measurement problem of this channel, it has now been quantified at scale, and almost every AI visibility dashboard in use today reports it as if it were signal.
How much actually changes
A study of 530,875 citations put daily URL-level churn between 44% and 88% depending on engine. Separate research tracking weekly stability found that once you step up from URLs to domains, 96.8% show no change week over week. Both numbers are true simultaneously, and the gap between them is the entire practical lesson.
The cause is not mysterious. These are probabilistic systems: sampling temperature, retrieval recency and minor prompt differences all move the output. Running the same twenty queries three times across five engines produces different citation sets each pass, and SparkToro's finding that AI recommendations change with nearly every query says the same thing from the brand-mention angle.
What this breaks
It breaks the daily screenshot. It breaks the single-run competitive audit. It breaks any report where someone asked the question once, got an answer they liked or hated, and drew a conclusion. Most damagingly, it breaks cause-and-effect reasoning: publish a change on Tuesday, see different citations on Wednesday, and the temptation to credit the change is enormous even though the same movement happens on days you ship nothing.
"If you did not measure a control, you did not measure anything. The citations move on the days you ship nothing too."
How to measure it properly
What to do this week
Fix your prompt set at ten questions, run each three times, record which domains appear, and put a date on it. Do that again next week. Resist the urge to look on Wednesday. In a month you will have four honest readings and a trend line, which is more than almost anyone in this field currently has, and it will save you from the much more expensive mistake of rebuilding your content strategy around a number that was noise.
