Method7 min readFielded

How to measure AI visibility: a question set, repeated runs, a coding scheme and honest margins

Measuring AI visibility is closer to running a small survey than checking a ranking. The questions have to be chosen carefully, each one asked more than once, every answer coded the same way, and the result reported with its uncertainty. This field report walks through that method step by step, with the extra care adult brands need.

Illustration of a bar chart with error whiskers on ivory, labeled Report 02

Q1Build a fixed question set

  • Collect the questions customers really ask: from site search, support emails, reviews and sales calls.
  • Group them by stage: discovery, comparison, purchase.
  • Write them in safe-for-work, natural language; avoid naming your brand in the question.
  • Freeze the set, and change it only on a schedule, so months stay comparable.

Q2Repeat every question

A single run tells you what one answer said, not what the assistant tends to say. Run each question several times on each assistant, in clean sessions, on similar days. More runs narrow the uncertainty; the audit page shows how quickly margins shrink as the sample grows.

Q3Code every answer the same way

CodeLabelRule
MMentionedBrand named anywhere in the answer
CCitedBrand’s own page given as a source
AAbsentAnswer given, brand not named
RRefusedAssistant declined or limited the answer on policy grounds
XInaccurateBrand named with a wrong or outdated detail

One answer can carry more than one code, for example M and X together.

Q4Keep refusals apart

Assistants sometimes limit answers on adult topics. Counting those as absences would make a lawful brand look invisible when the real cause is policy. Report the refusal rate on its own line, and compute visibility from answered questions only, stating which approach you used.

Q5Report the range, not a point

Give each figure with its sample size and margin, by assistant and question group. A move of a few points inside the margin is noise, not news. The service built on this method is the LLM visibility audit, and turning visibility into a comparative figure is covered in our share of voice guide.

Q6A coding sheet, column by column

ColumnWhat goes in it
Question ID and groupWhich question, and whether it is discovery, comparison, buying, trust or problem
Assistant, date, locationWhere and when the answer was produced
Full answer textSaved exactly, so anyone can check the coding
Brands named, in orderEvery brand mentioned and its position
Recommended brandThe one the answer suggests, if any
Cited linksEach source shown
Refusal or limitWhether the assistant declined or softened the answer
ErrorsAny factual mistake about your brand

Q7How much variation to expect

Run the same question ten times and the brand list may change several times. Studies of repeated prompts have found that identical lists are rare. That is not a flaw to hide; it is the nature of the tools. It means a single answer is an anecdote, a handful of answers is a hint, and only repeated samples describe what an assistant tends to say. Plan the number of runs before you start, based on how precise the figures need to be.

Q8Agreeing on the coding

Coding is a judgment. Is a brand "recommended" if it is listed first, or only if the answer says to choose it? Write the rules down, ask two people to code the same fifty answers independently, and compare. Where they disagree, refine the rule. Repeat until agreement is high, then keep the rules fixed. Without this step, a change in who does the coding can look like a change in the assistants.

Q9Storing the evidence

Save every answer in full, with the date, assistant, location and question, in a format you can search later. Months from now, someone will ask why a figure changed, and the answers are the only way to check. Store them securely, since they reveal your competitive interests, and keep them for at least as long as you report trends.

Q10A pilot round first

Before the first full round, run a small pilot: a dozen questions, a few runs each, on two assistants. It shows which questions produce useful answers, how often refusals happen in your category and whether the coding rules work. Adjust the set and rules, then start the real baseline. A pilot of a day or two saves weeks of figures built on a flawed set.

Q11Reports versus dashboards

Dashboards make it tempting to watch numbers that move by chance. A short written report each round, with figures, margins and a plain explanation of what changed, often serves decisions better. If you do use a dashboard, show margins on every chart and mark changes that exceed them. Share of voice calculations are covered in our share of voice guide.

Q12Choosing which assistants to measure

Measure the assistants your buyers actually use, not every assistant available. Your own analytics, customer surveys and the sources of referral traffic give clues. Most adult brands start with ChatGPT, Google's AI features, Perplexity and one or two others, then add or drop assistants as usage changes. Record the reasons for each choice, so later readers understand why the set looks the way it does.

Q13Writing the questions

Write questions the way buyers ask them, in full sentences or short searches, safe for work where buyers would phrase them that way. Avoid leading questions that mention your brand unless you are measuring brand-specific answers on purpose. Keep a separate set for questions that name your brand, such as "Is Brand X reliable?", since those measure reputation rather than discovery.

Q14Time and cost, roughly

TaskRough effort for 20 questions, 5 runs, 4 assistants (illustrative)
Running and saving 400 answersOne to two days, less with automation where terms allow
Coding 400 answersTwo to three days for one coder, plus checks
Analysis and reportOne to two days
Later roundsFaster, as rules and templates exist

Q15Linking visibility to outcomes

Visibility figures are most useful when set beside your own outcomes: branded searches, direct visits, sign-ups and orders. Track them on the same calendar as your rounds. If visibility rises and outcomes follow over several rounds, the link is worth investing in; if not, the measurement is still useful as an early warning, but not as a target. Rival analysis is covered in AI competitor analysis.

Q16Refusals in adult categories

Assistants refuse or soften some adult questions, and the rate varies by assistant, wording and time. Record each refusal, with the wording that triggered it, and report the refusal rate as its own figure. A rising refusal rate can explain a falling share better than any change in your brand's standing, and it tells you which questions to rephrase in the safe-for-work terms buyers also use.

Q17Next reads

The services built on this method are the LLM visibility audit, AI share of voice and competitive AI visibility. Sources are analysed in citation source analysis, and the terms behind it all in GEO vs AEO vs LLMO.

Q18One last check

Before sharing any figure, ask whether someone else could reproduce it from your saved answers and rules.

FAQQuestions

Can we do this ourselves?

Yes, on a small scale. The discipline lies in keeping the question set fixed, running each question several times, and coding answers consistently from month to month.

Should we log in when testing?

Use clean sessions without personal history where you can, and record the settings used, since personalization can change answers.

How do we treat an answer that refuses the question?

Record it as a refusal, not as an absence. Refusals are reported separately so they do not drag down a visibility figure unfairly.

LLM visibility audit

Find out how often AI assistants name your brand

Tell us your brand, your market and your main competitors. We run a sample of real customer questions through the major AI assistants several times, record mentions and citations, and send back a readout showing the ranges observed. The first one costs nothing and binds you to nothing.

Request an LLM visibility audit