Experiment
4 minute readOne AI Answer Is a Snapshot: Three Experiences with Citation Variability
People tracking AI answers report changing sources and unstable opportunity labels. Repeated, consistent sampling makes visibility claims more useful.
The lesson
Measure how often a brand or URL appears across repeated comparable answers, rather than treating one answer as a fixed ranking.
What people reported
u/-amphisbaena: Repeated queries returned changing sources
Reddit user u/-amphisbaena reports tracking 20 queries across ChatGPT, Gemini and Perplexity for nine days. They noticed different cited sources between runs and described differences across tools. Logged-in sessions and changes to the collection setup limit interpretation; the post does not prove a time-of-day mechanism or a general platform stability ranking. Read the original report: u/-amphisbaena — Reddit
u/EmbarrassedBuddy9743: A single pass misclassified content opportunities
In an r/aeo comment, u/EmbarrassedBuddy9743 says repeated work revealed prompts switching between apparently empty citation opportunities and strong incumbent sources. They describe running each prompt at least five times and assigning a majority category. This is a reported practical correction, not evidence that five runs are statistically sufficient for every question. Read the original report: u/EmbarrassedBuddy9743 — Reddit
Dana Billingsley: Initial inclusion did not always last
Dana Billingsley comments on Kim Huynh’s experiment that she has seen content appear quickly in AI answers and then disappear, while other sources were reused. She provides no sample counts or measurement schedule. Her qualitative experience is consistent with treating first inclusion as provisional, while leaving the reasons for retention unresolved. Read the original report: Kim Huynh — LinkedIn
What the experiences have in common
The shared problem is a changing observation being treated as a fixed property of a website. A team can find its brand once, celebrate an improvement and then conclude that a later absence means a penalty. Without repeated comparable measurements, either conclusion can be premature.
Create separate fields for a brand mention, an explicit recommendation and a linked source. A response can contain one without the others. Also separate a new conversation from an ongoing conversation, since the user’s preceding question changes the task. The measurement should describe the surface actually tested rather than implying it represents every customer conversation.
What these reports cannot establish
These sources are self-reports, including two Reddit pseudonyms and a short LinkedIn comment. No common prompt set or verified dataset is available. Changing retrieval results, model behavior, geography, account context and collection errors are possible explanations. A correlation with run time does not establish that time caused the change.
A test you can run: proposed protocol
- Fix a question set covering real buyer intents. Store exact wording, language, country context, product surface, visible model label and account state.
- Run each question repeatedly in comparable fresh sessions over at least two weeks, using permitted manual access or authorized collection. Keep follow-up conversations as a separate cohort.
- Save the answer, cited URLs, timestamp and any retrieval failure. Count errors separately instead of recording them as brand absence.
- Calculate mention and citation frequency per question, then summarize across questions with the same intent. Show the number of valid observations.
- Before crediting an intervention, compare its movement with unchanged questions and the baseline range. Repeat an apparent win before using it in a forecast.
The practical takeaway
A visibility report becomes more credible when it shows variation. The objective is a reproducible observation process, not a screenshot that happens to contain the brand.
Sources and research notes
Sources reviewed on September 15, 2026. Public social pages and search extracts sometimes expose inconsistent relative dates; unverified publication dates are omitted. Reported results are attributed claims, not independently audited facts. Reposts of the same underlying campaign are not counted as additional experiments.
- u/-amphisbaena — Reddit — Public post text retrieved.
- u/EmbarrassedBuddy9743 — Reddit — Comment read in parent thread; standalone comment fetch unavailable.
- Dana Billingsley — LinkedIn — Comment retrieved within Kim Huynh’s post.
Sources
- u/-amphisbaena — Reddit reddit.com
- u/EmbarrassedBuddy9743 — Reddit reddit.com
- Kim Huynh — LinkedIn linkedin.com
Community evidence review. These are attributed public reports, not experiments run by SEOVision. We did not access the participants’ analytics or independently reproduce their outcomes. The test below is a proposed protocol, with no SEOVision results claimed. Sources were reviewed on September 15, 2026; social posts may later be edited, removed or placed behind a login. Reported figures are attributed claims, not audited results.
Verification labels are shown only when a real review record exists. Demonstration content is not presented as independently tested.
Keep reading