LLM visibility audit
An LLM visibility audit that reports ranges, not hunches
An LLM visibility audit shows how often AI assistants name your brand when customers ask real questions. We build a question set from your market, run each question repeatedly on every assistant that matters, record what appears, and report visibility with the margin of error that the sample allows. For adult brands we also separate low visibility from policy limits on adult topics.
Visibility by assistant, repeated samples
- Assistant 144%
- Assistant 228%
- Assistant 362%
- Assistant 412%
What we record in every answer
- I-01MentionWas the brand named at all in the answer?
- I-02CitationWas the brand’s own site linked or cited as a source?
- I-03FramingWas the mention favorable, neutral, mixed or inaccurate?
- I-04CompanyWhich competitors appeared alongside it?
Across many answers, these records become the share-of-voice figures described on AI share of voice.
How sample size limits certainty
A visibility figure is only as firm as the number of answers behind it. At a true rate near 50 percent, the usual 95 percent margin of error shrinks like this:
| Answers sampled | Approximate margin |
|---|---|
| 30 answers | ± 17.9 points |
| 100 answers | ± 9.8 points |
| 400 answers | ± 4.9 points |
Standard formula for a proportion. Repeated AI answers are not perfectly independent, so we treat these margins as a floor, not a promise. The strategic use of the figures is covered on GEO strategy.
Not in our readouts
- Single screenshots presented as proof
- Figures without the sample size behind them
- Explicit test prompts designed to provoke adult answers
- Audits for escort or sexual-services brands
The coding scheme
| Code | Meaning |
|---|---|
| Named | Your brand appears in the answer |
| Recommended | The answer suggests choosing you |
| Cited | A link to your own site appears as a source |
| Accurate | What is said about you matches the facts |
| Sentiment | Positive, neutral or negative framing |
Two people code a sample of answers independently, and disagreements refine the rules, so results are consistent from round to round.
Why assistants differ
| Assistant type | What tends to drive its answers |
|---|---|
| Search-connected answers | Live web results and the pages it chooses to cite |
| Model knowledge without browsing | What the model learned in training, which can be out of date |
| Search engine AI features | The search engine's own index and ranking signals |
| Shopping-oriented answers | Product feeds, reviews and retailer listings |
Because each draws on different sources, a brand can be visible in one and absent in another. The audit reports each separately.
Accuracy errors we look for
- Products you no longer sell, or prices from years ago.
- Policies described wrongly, such as returns, delivery or discretion.
- Your brand confused with a similarly named one.
- Claims about safety or materials you have never made.
- Outdated legal or age-verification information.
Each error is traced to the source the assistant most likely used, so the fix starts in the right place rather than on your own site by default.
What the audit delivers
A written report with visibility, citation and accuracy figures by assistant and question group, each with its sample size; a list of the sources assistants cite in your category; a list of factual errors with likely origins; and a ranked set of fixes. The raw coded answers are included, so your team can check any figure. Repeat audits use the same question set and coding, so changes are comparable.
Location, accounts and personalisation
| Factor | How we control it |
|---|---|
| Location | Questions run from the countries and, where relevant, regions your customers are in |
| Signed-in versus signed-out | We record which; many runs are signed out to reduce personal history |
| Memory and past chats | Disabled or fresh sessions, so earlier questions do not colour later answers |
| Language | Questions asked in the languages your buyers use |
Personalization means no single run represents every user. Controlling these factors keeps rounds comparable, even if no setup can reproduce every buyer's experience.
When models change
Assistants update their models, browsing tools and policies often, sometimes without notice. A sudden shift in results may come from an update rather than from anything your brand did. We note the date and any announced model or policy change for every round, and when a major change lands, we re-run part of the question set quickly to separate the assistant's change from yours.
Wording matters
Small changes in how a question is phrased can change the answer. "Best vibrator for beginners" and "which vibrator should I buy first" may produce different brand lists. We include a few paraphrases of the most important questions and report whether results hold across them. Where a finding depends on a single wording, we say so plainly rather than presenting it as a general truth about the assistant.
Measuring responsibly
We run questions at a modest volume, through normal interfaces or official APIs, and within each provider's terms. We do not flood assistants with automated queries, try to manipulate their answers, or create fake reviews or posts to influence them. The aim is an honest picture of what buyers see, not an attempt to game it.
How long an audit takes
A first audit usually takes two to three weeks: one to agree the question set, competitors and coding rules, one to run and code the answers, and a few days to write the report. Repeat audits are noticeably faster because the question set and coding rules are already in place. We share early findings during the run, by email or a short call, if something urgent appears, such as a serious factual error about your products.
How many answers are enough
| Answers per assistant | What the figures can support (illustrative) |
|---|---|
| About 30 | A rough sense of whether you appear at all |
| About 100 | Shares precise to within roughly ten points |
| About 300 | Shares precise enough to compare rounds with confidence |
| More | Finer cuts by question group or market |
Rough guides based on standard sampling margins. Larger samples cost more time; we agree the size with you based on the decisions the audit needs to support.
After the audit
The report ends with a ranked list of actions, each linked to the evidence behind it: a page to improve, a third-party listing to correct, a question group to target. Some brands act on it themselves; others ask us to measure again after their changes. Either way, the next audit uses the same set and rules, so you can see what moved and what did not. Competitive follow-up is described on competitive AI visibility.
Who reads the report
Reports are written for founders and marketing leads, with a one-page summary at the top and the detail behind it. Developers and content teams get the specific pages and sources to act on, and every figure carries its sample size so nobody over-reads a small result or misses a large one.
Audit readout figures
| Figure | Typical use |
|---|---|
| Visibility by assistant | Where to focus first |
| Visibility by question group | Which buyer questions you are missing |
| Citation share | How often your site is a source |
| Error count | How much misinformation needs correcting |
Fees are on pricing, and prompt sets in our note on prompt set monitoring.
Questions
How many times do you ask each question?
Enough to report a range rather than a single figure, typically several runs per question on each assistant. The readout states the sample size behind every number.
Why might an AI assistant leave out an adult brand?
Sometimes because of weak visibility, and sometimes because the assistant limits answers on adult topics. We test neutral, safe-for-work phrasings to tell the two apart.
Will ChatGPT treat adult topics differently soon?
OpenAI announced an adult mode for verified adults in 2025, then delayed it more than once. We check each assistant’s current behavior at the time of every audit rather than assume.
What do we receive?
A written readout: visibility by assistant and question group, competitor comparison, cited sources, and a short list of priorities.
How do you make coding consistent?
Two people code a sample of answers independently using written rules, compare results and refine the rules until they agree. The same rules are used in every round.
LLM visibility audit
Find out how often AI assistants name your brand
Tell us your brand, your market and your main competitors. We run a sample of real customer questions through the major AI assistants several times, record mentions and citations, and send back a readout showing the ranges observed. The first one costs nothing and binds you to nothing.
Request an LLM visibility audit