Article
6 minute readBuild an AI Visibility Baseline You Can Reproduce
Create a reproducible AI visibility baseline by fixing prompts, settings and scoring rules while recording answer variability.
One AI answer is a snapshot, not a stable ranking. To learn whether visibility changes, keep the measurement conditions as consistent as the platform allows and preserve the underlying answers.
Define the measurement unit
Use one recorded answer to one exact prompt under documented conditions as the basic unit. Record platform, visible model or mode, date, locale, sign-in state and whether the conversation started fresh.
Some platform settings or underlying changes may not be observable. Mark those as unknown. “Same conditions” should mean the conditions you could control, not a claim that the provider's system stayed unchanged.
Establish a repeatable collection plan
Choose a fixed prompt panel and a manageable repeat schedule. For an initial operational pilot, three runs per prompt can reveal obvious variation, but this is a practical starting point rather than a statistical guarantee.
Save exact answers, citations, errors and timestamps. Follow platform access rules and use supported collection methods. Do not quietly rerun only the answers where your brand was absent.
Worked example: a misleading improvement
In a hypothetical panel of ten prompts, your brand appears in five answers on Monday and seven on Friday. It is tempting to report a 20-percentage-point improvement.
Repeated runs might show that the same prompts naturally vary between four and eight mentions. The two snapshots alone do not distinguish a durable change from ordinary variability. Report the observed counts and continue collection before assigning a cause.
Keep a baseline record
| Record | Why it matters |
| Exact prompt and version | Prevents wording changes from hiding in trends |
| Platform and mode | Separates different answer systems |
| Locale and session context | Documents relevant conditions |
| Raw answer and citations | Enables rescoring |
| Failure status | Preserves the denominator |
Separate measurement changes from market changes
Annotate provider updates, prompt revisions and scoring changes on the report. If you switch tools, overlap the old and new methods for a short period where feasible. Differences between tools may reflect coverage rather than a change in your brand.
Do not average platforms into a single score unless you explain the weighting and purpose. An equal-weight dashboard is a reporting choice, not a measure of market share.
Use the baseline to identify repeatable questions: which relevant prompts consistently omit the brand, which facts are repeatedly wrong, and which sources recur? Those patterns are more actionable than celebrating a single favorable answer.
Decide how much variation you can tolerate
Every answer system produces different output for the same prompt on different runs. That is not a flaw in your measurement; it is a property of the thing being measured, and the baseline has to describe it before it can describe anything else.
Run the panel several times over a few days before reporting a single number. For each prompt, note the range: in how many runs did the brand appear, and how did the recommendation vary? Prompts fall into three groups:
- Stable presence. The brand appears in nearly every run. Changes here are worth investigating.
- Stable absence. The brand never appears. These are the clearest opportunities and the easiest to track, because any appearance is a signal.
- Unstable. Appearance varies from run to run. For these prompts, a single snapshot is close to meaningless, and only a trend across many runs can be interpreted.
Report the group sizes with the baseline. A panel where half the prompts are unstable is telling you that week-to-week percentages will move on their own, and the report should say so before a stakeholder reads a change into noise.
What to freeze and what to allow to change
A baseline is only reproducible if the things that can be held constant are held constant, and the things that cannot are recorded. A short freeze list:
| Freeze | Record but cannot control |
| Exact prompt text and version | The provider's model updates and system behavior |
| Platform, mode, and locale setting | Any personalization the platform applies |
| Fresh session for each prompt | Retrieval sources the system consulted |
| Collection window (same days, same rough time) | Outages and partial failures, kept in the denominator |
| Scoring guide version | Reviewer judgment on borderline cases |
The right-hand column is why "same conditions" should always be described as "the conditions we control". When a provider announces a change, mark the date on every chart. A step change that coincides with an announced update is not evidence that your work did anything.
Turn the baseline into questions
The purpose of the baseline is not the headline percentage; it is the list of specific, repeatable observations that a team can act on. After the first full collection, produce three short lists:
- Prompts where the brand is consistently absent and a competitor is consistently present. For each, note what the cited sources cover that your site does not.
- Facts that are repeatedly wrong across runs and platforms, with the source cited when one exists. These become correction tasks with a clear owner.
- Sources that recur across many answers. These are worth reading closely, whether or not you can influence them, because they show what the answer system treats as authoritative for your topic.
Each list item is a hypothesis with a proposed check, not a conclusion. The baseline's value compounds when the second collection can be compared against these lists rather than against a single number that moved for reasons nobody can name.
Put this into practice
Copy the worksheet columns below into a spreadsheet and keep one row per item you check. The filled row is an illustrative example, not a reported customer result; replace it with your own verified records.
| Prompt ID | Version | Platform | Mode | Locale | Timestamp | Run status | Mention | Raw answer file |
| P01 | v1 | Specify platform | Specify mode | Specify locale | ISO timestamp | Valid | Pending | answers/R01.txt |
Use the following prompt only after supplying the records it requests:
Audit this AI visibility log for comparability. Identify changed prompts, modes, locales, missing runs and scoring rules. Summarize observed variation without attributing it to our content changes.Research context
Generated answers vary, making stable collection and scoring rules important for trend interpretation. The related Ahrefs starting points are You Can’t Track AI Like Traditional Search. Here’s What to Do Instead. and AI Overviews Change Every 2 Days (But Never Change Their Mind). This guide’s checklist, examples and proposed workflow are independently written; they are not results of a SEOVision experiment.
Continue with the next task
- Choose AI Search Prompts That Match Real Buyer Questions
- GEO vs SEO: What Changes in AI Search—and What Does Not
- AI Search Optimization: A Practical Guide Beyond GEO Hype
Sources
- You Can’t Track AI Like Traditional Search. Here’s What to Do Instead. — Research starting point; not an endorsement of this original workflow
- AI Overviews Change Every 2 Days (But Never Change Their Mind) — Research starting point; not an endorsement of this original workflow
Sources
- You Can’t Track AI Like Traditional Search. Here’s What to Do Instead. ahrefs.com
- AI Overviews Change Every 2 Days (But Never Change Their Mind) ahrefs.com
Examples are explicitly hypothetical and the workflow is an original SEOVision proposal, not a claimed experiment or a reported customer result. Sources were reviewed on September 15, 2026; platform behavior changes, so check the linked documentation before relying on any product detail. No ranking or traffic outcome is guaranteed.
These notes describe how this article was researched and what it does not claim. Guidance is educational; test any change on your own site and measure the result before relying on it.
Keep reading