How to Run Test Conversations for an AI Visibility Audit
Build a repeatable AI visibility testing workflow with versioned questions, recorded conditions, complete evidence, reviewer controls, and comparable reruns.
By Gaurav·Published ·Updated
Key Takeaways
- A repeatable AI visibility test needs a defined test record, not just a prompt and an answer.
- Keep a frozen benchmark question set for comparison and a separate exploratory set for new questions, competitors, and provider behavior.
- Record the provider surface, visible model information, locale, account state, retrieval state, date, and rerun count; mark unavailable conditions as unavailable rather than inferring them.
- Preserve the complete conversation, citations, retrieved sources where available, competitors surfaced, and the evidence behind every classified outcome.
- Compare cycles only after confirming the benchmark version, run conditions, outcome definitions, rerun counts, and reviewer calibration.
Introduction
By the time you reach this step, the temptation is to think the hard part is over. You've mapped your brand, built personas a real buyer would recognize, identified every scenario worth testing, and written questions that sound like something a person would actually type into a chat window. Running them feels like the easy part — paste the question in, read the answer, write down what happened.
It isn't, for the same reason a single search in ChatGPT was never an audit in the first place. AI models aren't deterministic — the same question asked twice can produce two different answers, on the same provider, in the same session. Run it on a different provider and the gap can be larger still. Treat any single exchange as your "result" for a scenario, and you're not measuring your visibility. You're measuring a sample of one, with all the same problems this series opened with — except now it's dressed up as a structured methodology, which makes it more convincing and just as wrong.
This post covers the complete measurement workflow: defining a stable test record, separating benchmark questions from exploratory questions, running each provider surface on its own terms, preserving the evidence, reviewing outcomes consistently, and deciding whether later cycles are comparable. It continues the previous posts on brand discovery, persona design, intent mapping, and question development.
The Measurement Workflow at a Glance
A repeatable AI visibility test has five stages:
- Set up the test. Define the buyer question, prompt version, provider surface, run context, and outcomes the test can observe.
- Run the test. Use the approved question or variant, preserve the provider-specific experience, and continue the conversation only where the scenario and surface support it.
- Record the evidence. Keep the complete response, visible citations, retrieved sources where available, competitors surfaced, and the conditions under which the answer was produced.
- Review the outcome. Apply shared definitions to observable evidence, add reviewer annotations, and leave ambiguous distinctions unresolved.
- Compare cycles. Confirm that benchmark versions, surfaces, run conditions, rerun counts, and review rules are comparable before interpreting change.
The value of the workflow is the evidence path it preserves. A reader should be able to move from a reported finding back to the provider response, the buyer question, the run conditions, and the rule used to classify the result.
Why it matters: Without that path, a trend line can look precise while silently mixing different questions, surfaces, or review rules.
Define One Test Record
One test record is the smallest inspectable unit of the workflow. It contains:
- a stable buyer question and its prompt version;
- the persona, buyer intent, and question context;
- the provider and product surface;
- the recorded run conditions;
- the complete response or conversation;
- visible citations and retrieved-source evidence where the surface exposes it;
- brand, competitor, recommendation, and source observations;
- the classified outcome; and
- the reviewer annotation that explains how the evidence met the definition.
The classified outcome does not need to be a numerical score. It is a documented interpretation of observable response evidence. If the evidence does not support a distinction, record the field as unresolved instead of forcing a classification.
Why it matters: A result is reusable only when another reviewer can inspect the same record and understand how the conclusion was reached.
Maintain a Versioned Prompt Library
Separate the prompt library into two sets.
Frozen benchmark set. Use this set for comparisons over time. Each entry should have a stable prompt identifier, version, effective date, persona, intent, context type, core buyer question, approved variants, conversation goal, and stop condition. A material question change creates a new version or starts a new comparison series.
Exploratory set. Use this set for newly observed buyer questions, competitors, product capabilities, provider behavior, and hypotheses that are not yet part of the benchmark. Exploratory results can inform future testing, but they should not be silently added to a benchmark trend.
Record why a prompt changed and when the new version became effective. Do not overwrite the prior wording in a way that makes older records impossible to reconstruct.
Why it matters: A stable benchmark supports comparison. A separate exploratory set lets the audit learn without quietly changing what the trend claims to measure.
Open With a Variant, Not Always the Core Question
Each scenario in your question set came out of Step 4 with a core question and two or three variants — different phrasings of the same underlying scenario. When it's time to run a scenario, don't reach for the core question every time. Rotate through the variants instead, so that across however many times you run a given scenario, no single phrasing is solely responsible for the pattern you see.
This matters more than it sounds like it should. Two buyers with the same problem, asking in slightly different words, can get meaningfully different AI responses — different sources pulled in, different brands named, different specificity in the answer. If every run of a scenario opens with the exact same sentence, you can't tell whether a result reflects how the AI treats that scenario or how it happened to treat that one sentence. Drawing from the pool of variants spreads that risk across the set instead of concentrating it in one phrasing.
Why it matters: The variants you wrote in Step 4 only do their job if they're actually used. A question set with three solid variants per scenario, run with the same one every time, has quietly thrown away the protection it was designed to provide.
Run the Conversation Forward, Not Just the Opening Line
A single exchange — one question, one answer — tells you whether your brand showed up to that exact question. It tells you very little about what happens once the buyer reacts to what they just read, which is what real buyers do. Real conversations have a second message: a narrowing question, a "what about," a request to compare two things the AI just mentioned.
Once the opening question gets an answer, write the next message the way the persona actually would — based on what the AI said, in service of the conversation goal you defined for this scenario back in Step 4. If the AI gave a category-level answer and the persona's goal was to land on something concrete, the natural follow-up narrows toward their actual constraints: team size, budget, timeline, the thing they specifically care about. If the AI named a competitor the persona would recognize, the natural follow-up might ask how your brand stacks up against it. The follow-up isn't scripted in advance — it responds to the conversation as it's actually unfolding, the same way a buyer would.
This is where a meaningful share of organic visibility actually shows up. A first answer that only lists category-level options can turn into a specific, source-backed recommendation two messages later, once the buyer has narrowed the conversation enough for the AI to commit to a name. A test that stops after one exchange would have recorded that first answer and missed everything that came after it.
Why it matters: Buyers don't ask once and walk away. A test that does is testing a different, easier-to-measure thing than the conversational visibility this whole process exists to capture.
Decide When a Conversation Is Done — and Stop There
Every scenario from Step 4 has a stop condition: the specific point at which the conversation has accomplished what it set out to do. Use it. After each response, check it against that condition. If it's been met — the buyer has a usable shortlist, understands the tradeoffs, knows whether the brand in frame fits — end the conversation there, even if there's more that could theoretically be asked.
Pair this with a hard cap on the number of turns, as a safeguard rather than a target. Most conversations will reach their stop condition well before the cap. The cap exists for the conversations that don't — where the AI keeps offering to go deeper, or the back-and-forth drifts without resolving, and nothing forces it to end on its own.
Why it matters: A scenario tested for one exchange in one run and five exchanges in another isn't being tested consistently. The stop condition and the turn cap together keep every run of the same scenario roughly comparable, so a difference in outcome reflects the AI's behavior rather than how long the conversation happened to continue.
Build for Repeatable Execution
Question variants, response-based follow-ups, stop conditions, and evidence capture need to be applied consistently. As the number of personas, scenarios, providers, and reruns grows, manual execution becomes harder to keep consistent.
Automation can help with repeatable submission, evidence capture, identifiers, timestamps, and storage. Human review is still required where the next buyer question depends on meaning, where an outcome is ambiguous, or where the provider surface does not expose a clean structured record.
The appropriate level of automation depends on the test scope and the surface being measured. The methodology does not require a universal number of conversations before automation becomes worthwhile. It requires a process that can preserve the defined conditions and evidence without silently changing how records are produced.
Why it matters: Repeatability comes from applying the same documented process and preserving the same kinds of evidence, not from reaching a particular volume or removing human judgment entirely.
Choose the Surface Before You Choose the Automation
The surface being tested is part of the result. A consumer chat product, a provider API, a conversational search product, and a search-results answer block may expose different models, retrieval behavior, citations, personalization, and interaction patterns.
Consumer conversational product. This measures the interface a buyer can use directly. Record the product surface, account state, locale, visible model information, and whether live retrieval appears to be enabled when those details are observable.
Provider API. This supports programmatic execution and structured response capture, but it is not automatically equivalent to the provider's consumer product. Record the model, tools, retrieval configuration, system instructions, and other exposed settings. Describe the result as an API result unless comparability with another surface has been established.
Conversational search surface. Google AI Mode can support follow-up questions, but it remains a Google Search surface and should be recorded separately from Gemini and from Google AI Overview.
Search-answer surface. Google AI Overview and Microsoft Bing AI Answers (Copilot Search) may show one generated answer or no answer block. Record whether an answer was produced. If there is no answer to evaluate, classify the record as excluded or unevaluated rather than treating the missing answer as a brand failure.
Choose the execution method that preserves the surface the test is meant to observe. Do not combine API, consumer-product, and search-interface results into one line unless the methodology establishes how they can be compared.
Why it matters: A well-automated test of the wrong surface produces a consistent answer to a different question.
Test Every Provider and Surface on Its Own Terms
Run the applicable scenarios separately across ChatGPT, Claude, Gemini, Google AI Mode, Google AI Overview, Grok, Microsoft Bing AI Answers (Copilot Search), and Perplexity. Do not run a scenario on one surface and assume the result holds for the others. See our experience here with Claude.
ChatGPT, Claude, Gemini, Google AI Mode, Grok, and Perplexity can support conversational follow-ups where the buyer scenario warrants them. Google AI Overview and Microsoft Bing AI Answers are search-answer surfaces: preserve the generated answer and sources when present, and record when no answer block appears.
Providers and surfaces can differ in retrieval, source selection, visible citations, model exposure, personalization, and willingness to name or recommend a brand. A brand can receive substantive, cited treatment on one surface and no answer reference on another under otherwise similar buyer questions.
Record the provider and product surface separately. A provider-level finding should not be generalized into a brand-wide finding unless the pattern is supported across the relevant surfaces.
Why it matters: Provider and surface differences are part of the evidence, not noise to average away.
Record the Run Conditions and Preserve the Complete Evidence
Store every test as structured data connected to the complete underlying evidence.
For each test record, preserve the following setup and run fields:
- record identifier;
- prompt identifier, version, and effective date;
- benchmark or exploratory designation;
- persona, buyer intent, and question context;
- provider and product surface;
- consumer, API, conversational-search, or search-answer interface;
- visible model or version information;
- locale or market;
- account and personalization state where observable;
- retrieval or web-search state where observable;
- test date and time;
- run identifier; and
- rerun count or repetition index.
When a surface does not expose a field, record unavailable. Do not infer the model version, retrieval state, locale, or account condition from the answer alone.
Preserve the response evidence:
- the opening question;
- every follow-up question;
- every provider response;
- the complete answer text rather than a summary alone;
- visible citation labels, domains, and URLs;
- retrieved sources where the surface exposes them;
- competitors and alternative approaches named;
- the turn on which the brand first appears;
- observable recommendation, comparison, or caveat language; and
- the evidence used to support each classified outcome.
Store the records so they can be filtered by persona, intent, context, provider, surface, prompt version, and run. A narrative summary may accompany the record, but it must not replace the answer and source evidence.
Why it matters: Conditions that were never recorded cannot be reconstructed reliably after the provider, interface, or prompt library changes.
Repeat Enough to Evaluate a Pattern
One answer is one observation. A pattern requires repeated evidence across the personas, intents, question contexts, providers, and surfaces relevant to the decision.
There is no universal number of runs that makes every finding reliable. The appropriate depth depends on the expected variability, the breadth of the question set, the consequence of the decision, and whether the observation recurs in the slices where it should appear.
Preserve the rerun count for every record. When a conclusion rests on one response, label it as an isolated observation. When it recurs, describe the provider, surface, persona, intent, and context boundaries of the pattern rather than generalizing beyond the tested evidence.
Why it matters: Repetition can show that an outcome recurs under defined conditions. It does not reveal the provider's internal cause or turn a narrow result into a universal one.
Review Outcomes Consistently
Apply written definitions to observable response evidence. At minimum, record separately:
- whether the provider produced an answer that could be evaluated;
- whether the brand appears in the answer;
- whether the buyer introduced the brand or the provider surfaced it independently;
- the primary evidence state: answer reference, mention-only treatment, source-only presence, no answer reference, or unresolved;
- whether an evaluated answer reference meets the qualified-visibility criteria;
- whether the brand is recommended, merely listed, criticized, or described conditionally;
- whether a source is visibly cited;
- whether a brand-owned source is visibly cited;
- which competitors or alternatives appear;
- whether a competitor receives materially stronger treatment than the target brand; and
- the reviewer annotation supporting each classification.
Do not collapse these observations into one unexplained score. When the evidence is incomplete, ambiguous, or outside the written definitions, record the distinction as unresolved.
Why it matters: Shared definitions make the outcome inspectable. They do not remove ambiguity, but they prevent ambiguity from being hidden inside a confident label.
Calibrate Reviewers Before Comparing Results
Before reviewed outcomes are treated as comparable, have reviewers apply the same definitions to a shared sample of records. Compare the classifications, discuss borderline cases, record the agreed interpretation, and add examples to the review guidance when a recurring edge case appears.
If a definition changes materially, record the methodology version and determine whether earlier records need to be reviewed again. Do not invent a universal reviewer-agreement threshold unless the methodology formally adopts one.
Why it matters: A stable prompt set does not produce a comparable trend when the meaning of the outcome changes between reviewers or cycles.
Compare Testing Cycles Under Documented Conditions
Run another cycle when meaningful implementation work, a provider change, a market change, or a defined decision justifies a new observation. Do not prescribe a universal weekly or monthly schedule.
Before comparing two cycles, confirm:
- the frozen benchmark version;
- the applicable prompt and variant versions;
- the persona, intent, and question-context definitions;
- the provider and product surfaces;
- the recorded model, locale, account, and retrieval conditions;
- the rerun counts;
- the outcome definitions and methodology version; and
- completed reviewer calibration.
Keep exploratory questions outside the benchmark trend unless they are deliberately promoted into a new benchmark version. When a prompt, surface, provider condition, or review rule changed, preserve the result and describe the change as comparison context. Label the comparison as limited rather than presenting unlike records as equivalent.
Why it matters: Trend interpretation is credible only when the team can see what stayed fixed, what changed, and how the change limits the conclusion.
What This Looks Like in Practice
The historical Freshdesk examples below show how provider behavior and buyer-led follow-ups can change the observed outcome for one scenario. They illustrate the evidence a complete conversation can preserve; they do not prescribe a required provider set, run count, or universal result.
SCENARIO: Early-Stage, No Category Frame
Question run: "Our support conversations are scattered everywhere, and I
can't see what's happening clearly. How should we think about fixing
that?"
ChatGPT — single exchange
Gives a structured framework for centralizing intake, standardizing
workflow, and building reporting. No platform named anywhere in the
answer.
Result: Freshdesk did not appear.
Claude — single exchange
Gives a similar framework, then names the category of tool that solves
it — "modern help desks like Zendesk, Intercom, and Freshdesk," all
described as unifying conversations, tickets, AI, and reporting in one
place.
Result: Freshdesk appeared — named in a peer list alongside four
competitors, not cited as a source, not recommended on its own.
Same persona, same question, same context level. Two providers, two
different outcomes — exactly why each one gets tested on its own terms.
---
SCENARIO: Early-Stage, No Category Frame — Seeking a Recommendation
Question run: "What kind of tool or approach helps bring conversations,
tickets, reporting, and AI together without making support harder to
run?"
Turn 1 (ChatGPT): Describes the category of solution, then mentions
modern help desks "like Zendesk, Intercom, and Freshdesk" — Freshdesk
appears, but only inside a peer list. A single-exchange test would have
recorded this and stopped here.
Follow-up (in the buyer's voice, narrowing toward their actual
situation): "Can you narrow that down to the best fit for a mid-sized
support team that wants simple setup and low admin overhead?"
Turn 2: Freshdesk becomes a standalone recommendation, supported by two
cited Freshworks pages, with Intercom named as the main alternative.
Follow-up: "Can you give me a simple Freshdesk vs. Intercom
recommendation for a mid-sized team that wants the easiest setup and
lowest admin overhead?"
Turn 3: Freshdesk recommended again, head-to-head against Intercom, with
specific reasoning about setup complexity and admin overhead.
A passing mention in turn one became a cited, head-to-head recommendation
by turn three — driven entirely by follow-ups written from the buyer's
actual goal, not a script. That shift is the reason this step asks for
conversations rather than one-off questions.
Within that recorded Freshdesk test set, organic appearance was weaker in fully unbranded conversations and stronger after the buyer introduced category, competitor, or brand context. That was an observation from the defined questions, providers, and conditions in that run—not a benchmark that another brand should be expected to reproduce. The analysis step should preserve those boundaries while investigating what the pattern may support doing next.
What You Have at the End of This Step
A completed testing workflow produces five connected outputs.
A versioned prompt library. The frozen benchmark set supports comparison, while the exploratory set captures new questions without silently changing the benchmark.
Recorded run conditions. Each result identifies the provider surface, visible model information, locale, account state, retrieval state, date, and rerun count, with unavailable values preserved honestly.
Complete response and source evidence. The record retains the full conversation or answer, visible citations, retrieved sources where available, competitors surfaced, and the passages behind each classification.
Reviewed outcomes. Shared definitions and reviewer annotations connect the evidence to appearance, answer reference, qualified visibility, recommendation, citation, and competitor observations without collapsing them into one unexplained score.
Comparison boundaries. A later cycle can show what remained comparable, what changed, and where interpretation must be limited.
This is the evidence the analysis step turns into discovery, framing, displacement, citation, provider, and trend findings. For how all six audit steps connect, see the complete overview. For the next step, continue to How to Turn AI Visibility Test Data Into Actionable Findings.
If you'd rather see what your brand's test results look like before running this yourself, fill out the form below.
Find out what ChatGPT says about your brand.
Share your website. We’ll test real buyer questions specific to your brand in ChatGPT and deliver a reviewed analysis showing where your brand appears, how it is framed, which competitors and sources shape the answers, and what to do next.
Full Viziquo analyses include ChatGPT, Claude, Gemini, Google AI Mode, Google AI Overview, Grok, Microsoft Bing AI Answers (Copilot Search), and Perplexity.
Free · No credit card needed · Delivered within 3 business days
