AEDaily SUBSCRIBE
Home / Guest Posts / The volatility data
GUEST POST · RESEARCH Aug 9, 2026 · 6 min read

Ask the same question twice, get different sources. The volatility data

PN Priya Natarajan Independent researcher · 6:00 AM ET 𝕏 in ✉
Editorial illustration: a calendar-like grid of repeated marks drifting in pattern over time

Ask an AI engine the same question twice and you will often get two different sets of sources. That is not a bug you are catching, and it is not your competitors outmanoeuvring you overnight. It is the defining measurement problem of this channel, it has now been quantified at scale, and almost every AI visibility dashboard in use today reports it as if it were signal.

⚡TL;DRStudies of hundreds of thousands of citations find engines replacing 44% to 88% of cited URLs from one day to the next, while the underlying domains stay about 97% stable week to week. Individual URL tracking is mostly noise. Domain level, aggregated across repeated runs, weekly, is the only reading that means anything.

How much actually changes

A study of 530,875 citations put daily URL-level churn between 44% and 88% depending on engine. Separate research tracking weekly stability found that once you step up from URLs to domains, 96.8% show no change week over week. Both numbers are true simultaneously, and the gap between them is the entire practical lesson.

What changes, and what does notURLs replaced, day to day44% to 88%Domains unchanged, week to week96.8%URL churn from a 530,875-citation study; domain stability from weekly stability research.

The cause is not mysterious. These are probabilistic systems: sampling temperature, retrieval recency and minor prompt differences all move the output. Running the same twenty queries three times across five engines produces different citation sets each pass, and SparkToro's finding that AI recommendations change with nearly every query says the same thing from the brand-mention angle.

What this breaks

It breaks the daily screenshot. It breaks the single-run competitive audit. It breaks any report where someone asked the question once, got an answer they liked or hated, and drew a conclusion. Most damagingly, it breaks cause-and-effect reasoning: publish a change on Tuesday, see different citations on Wednesday, and the temptation to credit the change is enormous even though the same movement happens on days you ship nothing.

"If you did not measure a control, you did not measure anything. The citations move on the days you ship nothing too."

How to measure it properly

PRACTICEWHY
Sample each prompt several timesOne run is a coin toss. Aggregate the runs before you record anything
Report at domain levelURLs churn 44 to 88% daily; domains hold at about 97% weekly
Use a weekly cadenceDaily readings are dominated by noise. Weekly is the shortest honest interval
Keep a fixed prompt setChanging the prompts between readings makes the series meaningless
Watch trend, not positionDirection over several weeks carries the signal; a single reading does not
Measurement practice implied by the volatility research.

What to do this week

Fix your prompt set at ten questions, run each three times, record which domains appear, and put a date on it. Do that again next week. Resist the urge to look on Wednesday. In a month you will have four honest readings and a trend line, which is more than almost anyone in this field currently has, and it will save you from the much more expensive mistake of rebuilding your content strategy around a number that was noise.

KEY TAKEAWAYS01Engines replace 44% to 88% of cited URLs day to day. That churn is the system working as designed, not a signal.02Domains are stable: about 96.8% unchanged week over week. Report at domain level.03One run is a coin toss. Sample every prompt several times and aggregate before recording.04Weekly is the shortest honest cadence. Daily readings mostly measure sampling noise.05Without a control period you cannot attribute a citation change to anything you shipped.