A visibility result is only useful when you can inspect how it was produced and understand whether a later run measured the same thing.
Version the questions and test rules
Each run snapshots the buyer personas, scenarios, question variants, and brand-introduction rules used in that test set. For conversational tests, the snapshot also includes the conversation goal, follow-up rules, and stopping conditions. We preserve the exact questions used in each run, including any response-dependent follow-ups.
Preserve the evidence behind each result
Each test record connects:
The test context: run ID, timestamp, persona, scenario, intent, context level, and the prompt and taxonomy versions for that run.
The execution conditions: provider and model identifier when disclosed, plus search and citation evidence the provider exposes.
The response evidence: the exact questions and answers, preserved turn by turn, with exposed source URLs and citations.
The observed outcomes: brand appearances, recommendations, competitor appearances, and source ownership. Brand appearance and retrieval include the turn on which they first occurred.
Where a value is not observed, it is left unset rather than inferred.
Keep different visibility signals separate
Retrieved means a page appears in retrieval evidence exposed by the provider. Cited means the answer explicitly links to or attributes a source. Mentioned means the brand is named. Recommended means the answer presents the brand as a suggested choice or fit.
These observations can overlap, but they are not interchangeable. A citation does not automatically mean a recommendation, and an uncited page cannot be assumed absent from retrieval when retrieval evidence is unavailable.
We also distinguish organic from prompted appearances by recording whether the buyer had introduced the target brand before the answer being evaluated.
Rerun comparable tests—and explain what changed
Reruns follow an agreed review cadence or respond to meaningful content, provider, or competitive changes. For comparison, we retain the scenario set, question-variant pool and sampling approach, follow-up rules, and outcome definitions, while keeping execution conditions consistent where possible.
Material changes are versioned as a new run. Results retain their sample counts. Repeated tests can produce different answers, and a before-and-after change alone does not prove that a website edit caused it.