How to Measure AI Visibility: Mentions, Citations, Share of Voice and a ChatGPT Brand Test

Everyone wants to know their AI visibility, yet most teams still judge it by gut feeling. Someone asks ChatGPT about the brand, sees the name in the answer and moves on. However, ask the same question five times and you may get five different lists.
I have worked in search marketing since 2012. In this guide I share the protocol I use to measure AI visibility in a repeatable way: a prompt set, platform sampling, six metrics, a tracking sheet, a monthly report and a 10 minute ChatGPT test. I also covered how to grow visibility in a separate guide on how your brand shows up in ChatGPT and Gemini. This one, however, is only about measurement.
How do you measure AI visibility?
AI visibility is how often ChatGPT, Gemini, Google AI Overviews, AI Mode, Perplexity and Copilot mention your brand, cite your site and describe you correctly. To measure it, you run a fixed prompt set several times on each platform, log every answer in a sheet and compare your rates with those of your competitors.
The key point is simple: one screenshot is not a measurement. AI answers shift a little every time, so any conclusion based on a single run can mislead you. Measurement means asking the same question under the same conditions many times and talking in rates. In short, AI visibility is a probability, not a rank.
The protocol in this guide has four layers: the prompt set, platform sampling, metrics and reporting. I cover each layer under its own heading, so you can jump to the part you need. At the end you will also find a 10 minute manual test and a monthly report template.
Why don't classic SEO reports answer this question?
Classic reports show where your page ranks and how many clicks it earns. An AI answer, on the other hand, has no list of ten blue links; it has a paragraph, a short list or a few source cards. In practice, users often decide without clicking at all. As a result, click data tells only a small part of the story.
Official reports now close part of this gap. Google shows AI Overviews and AI Mode impressions in a separate Search Console report. Microsoft also counts citations in Copilot answers inside Bing Webmaster Tools. I explain both sources in detail further down, because each has clear limits.
Still, two gaps remain. First, neither report covers ChatGPT, the Gemini app or Perplexity. Second, neither one counts answers that name your brand without a link to you. In other words, if an AI recommends you without a link, the official reports show nothing. That is why you need your own way to measure AI visibility.
What is the difference between a mention, a citation and share of voice?
People mix up these three terms all the time, yet each one answers a different question. A mention means your brand name appears in the answer text. A citation, on the other hand, means the answer links to a page on your site as a source. Share of voice is your mentions as a share of all mentions of you and your competitors across the same prompts.
The difference matters in practice. For example, an answer can recommend you and still cite a competitor's blog post; that is a mention without a citation. The reverse also happens: the answer cites your guide but never names your brand. The first signals brand awareness, while the second signals content authority.
Share of voice also needs context. When the competitor list changes, the rate changes too, so fix your competitor set at the start and keep it stable within the month. Also calculate every metric from the same set of answers. If you mix answers from different days, the rates stop explaining each other.
Which six metrics should you use to measure AI visibility?
I build every measurement on six metrics. First I mark each one at answer level, then I turn it into a rate across the whole prompt set:
- Mention rate: the number of answers that name your brand divided by all answers.
- Citation rate: the share of answers that cite at least one page from your site.
- List position: where you sit when an answer gives a list; track your top three rate rather than an average position.
- Share of voice: your mentions divided by the total mentions of you and your competitors.
- Sentiment: whether the answer describes your brand in a positive, neutral or negative tone.
- Accuracy: how well the facts about your brand, such as services, location, pricing policy and founding year, match reality.
Above all, these six metrics work as a set. For example, a high mention rate with low accuracy shows that the AI knows you but describes you wrongly. A high citation rate with a low mention rate suggests that AI tools use your content while your brand does not stick. So never report any of them in isolation.
How do you design a prompt set?
A prompt set is the list of questions you repeat word for word every month. Its strength, however, comes from balance rather than size. I combine four prompt types, because each one stands for a different customer moment:
- Brand prompts: “What does [Brand] do?” and other questions that include your name.
- Category prompts: “Which web design agencies in Manchester are best for B2B firms?” and other discovery questions without your name.
- Comparison prompts: “[Brand] or [Competitor]?” and “What are the best alternatives to [Competitor]?” at the decision stage.
- Problem prompts: “Why doesn't my site show up on Google?” and similar questions where your solution could come up.
Thirty prompts are enough to start: 5 brand, 12 category, 6 comparison and 7 problem prompts. The weight on category prompts is deliberate, because the questions that bring new customers rarely include your name. Also write every prompt in the words your customers actually use. The phrases your sales team hears on the phone make better prompts than any keyword list.
How do you write brand and category prompts?
Brand prompts show whether the AI knows you and how it describes you. Keep them plain. For example, ask “Which services does [Brand] offer?” and note what comes back. That way you read accuracy and sentiment from the same answer.
Category prompts are where the real competition happens. A user asks for “the best accounting software for a small business” without knowing your name, and the answer offers a handful of options. If you are not on that list, success on brand prompts brings no new customers. So add qualifiers such as location, budget and industry, because real questions almost always carry them.
Give every prompt a code, such as B01, C01, V01 or P01. The code then lets you compare the same question month after month. If you change a prompt, give it a new code; otherwise old and new results blend together and the trend line loses its meaning.
Why do comparison and problem prompts matter?
Comparison prompts stand for the last step before a purchase. When someone asks “What are the differences between [Brand] and [Competitor]?”, the AI places both brands side by side and often leans towards one of them. In these answers, sentiment and accuracy tell you more than the mention rate.
Problem prompts, however, sit at the top of the funnel. The user is not looking for a vendor yet; they are describing a pain. For instance, someone who writes “too many shoppers abandon the cart on my store” may get a service or tool suggestion in the answer. If your brand shows up there, the AI links you to that problem.
Without these two types, your measurement only shows brand awareness. Yet growth depends on appearing in the questions of people who do not know your name yet. Therefore, reserve at least a third of the set for comparison and problem prompts.
Which AI platforms should you sample?
Start with the platforms your audience actually uses; you do not have to measure all of them with the same intensity. Specifically, the table below sums up what I log on each platform and what I watch out for.
| Platform | How to sample | What to watch |
|---|---|---|
| ChatGPT | Unpersonalized Temporary Chat | Note whether the answer came from a web search or from the model's own knowledge |
| Gemini | Temporary chat | Log in a separate column whether the answer links to sources |
| Google AI Overviews | Signed out browser, target country and language | Not every query triggers one; record whether an overview appeared |
| Google AI Mode | Same query in the AI Mode tab | Record the first answer without follow up questions |
| Perplexity | Clean account, default settings | Record which URLs sit behind the numbered sources |
| Copilot | Clean account, default settings | Compare results with Bing Webmaster Tools data |
For small and mid sized brands I start with ChatGPT, AI Overviews and Gemini. I add Perplexity and Copilot when the audience is technical or corporate. If you wonder which assistant is more reliable for which kind of question, see my comparison of ChatGPT, Gemini, Claude, Copilot and Perplexity.
Why do AI answers change every time?
Language models generate answers word by word, based on probabilities. The same prompt can start with a different word on another run and drift in another direction. Systems that search the web add one more factor: the sources they retrieve in the background also change from run to run.
The scale of this variation is not small. In research that SparkToro and Gumshoe published in January 2026, 600 volunteers ran 12 prompts through ChatGPT, Claude and Google's AI answers a combined 2,961 times. According to the study, the chance that ChatGPT or Google's AI returned the same list of brands twice was below 1 in 100. Moreover, the chance of the same list in the same order was roughly 1 in 1,000.
This has two consequences. First, a claim like “we rank second in ChatGPT” means nothing without repeated runs. Second, the meaningful metric is frequency rather than position: in how many answers your brand appears. In fact, the researchers themselves recommend tracking a visibility percentage instead of a rank. I also report position only as supporting information.
How many times should you run each prompt?
One run is not a measurement, yet unlimited runs do not fit any budget. My practical rule: run each prompt at least 3 times per platform, and your critical prompts 5 to 10 times. With 30 prompts and 3 runs you collect 90 answers on one platform; across three platforms that becomes 270.
Sample size decides how much your rate can move by chance. A simple statistical estimate puts the 95 percent confidence interval for a rate near 50 percent at roughly ±10 points with 90 answers. With 270 answers, however, it shrinks to roughly ±6 points. So in a 90 answer sample, a rise from 40 to 46 percent may be pure noise.
That is why I apply two rules to monthly changes. First, I treat any change smaller than the margin of error as flat. Then I accept a change as real only when it moves in the same direction two months in a row. As a result, I never report a random spike as a win. Put simply, when you measure AI visibility with small samples, only large moves mean something.
How do you keep the setup stable when you measure AI visibility?
To make your results repeatable, you need the same conditions every month. Otherwise you cannot tell whether a change came from your work or from the environment. These are the rules I follow:
- In ChatGPT I open a Temporary Chat and choose the Unpersonalized option before it starts; according to OpenAI's help page, that option does not use memory or custom instructions.
- In Gemini I also use a temporary chat, because in these chats Gemini does not tailor answers to your past conversations.
- On Google I use a signed out browser window and keep the language and country of the target market.
- Next to each record I note the date, platform, mode and, where visible, the model name.
- Finally, I finish each round in the same week every month, ideally within two days.
If you sell in several markets, keep a separate set for each one. English and German prompts draw on different sources, so never pool their results in one table. Likewise, if you care about mobile versus desktop differences, track them in a separate column. For UK and US audiences, run separate rounds too, since local providers and sources differ.
What should your tracking sheet look like?
In the tracking sheet, every row is one answer. That way the three runs of a prompt sit in three rows, and you calculate the rates directly. These are the columns I use to measure AI visibility:
| Column | What you enter | Example value |
|---|---|---|
| Date | Day of the round | 29 Sep |
| Platform and mode | Platform, search status, model | ChatGPT, web search used |
| Prompt code | Fixed code from the set | C03 |
| Prompt type | Brand, category, comparison, problem | Category |
| Run | Which attempt of the same prompt | 2 |
| Mention | Did the brand name appear | 1 |
| Position | Place in the list, blank if no list | 3 |
| Citation | Did the answer cite your site | 0 |
| Source URLs | Links shown in the answer | competitor.com/guide |
| Competitors named | Brands from your competitor set | Competitor A, Competitor C |
| Sentiment | Positive, neutral, negative | Neutral |
| Accuracy | Correct, incomplete, wrong | Incomplete |
| Notes | Anything unusual | Outdated pricing |
You can keep the sheet in Google Sheets or Excel. Above all, the columns must stay the same every month: adding a column is fine, removing one is not. Columns marked with 0 and 1 also make formulas easy. For example, the AVERAGE of the Mention column is your mention rate.
How do you calculate the metrics from the sheet?
Once the sheet is ready, the calculations are simple. Here is a worked example: you ran 30 prompts 3 times each in ChatGPT and collected 90 answers.
- Your brand appears in 27 answers, so your mention rate is 27/90, or 30 percent.
- Nine answers cite your site, which gives a citation rate of 10 percent.
- In 12 of those 27 answers you sit in the top three, so your top three rate is 12/27, or about 44 percent.
- Three of the 27 answers contain a factual error, which leaves an accuracy rate of 24/27, or about 89 percent.
These figures are only a worked example, not the result of a real brand. I calculate every rate both in total and by prompt type. The reason is simple: a 90 percent mention rate on brand prompts can easily hide a 5 percent rate on category prompts. Your total can look healthy while you are missing from the very questions that bring new customers. A conditional average on the prompt type column gives you that split in one formula.
How do you calculate share of voice against competitors?
For share of voice, first define a fixed set of 3 to 5 competitors. Then count which brands appear in each answer. The formula is your mentions divided by the total mentions of you and your competitors.
For example, if you appear 27 times in 90 answers and your three competitors appear 63 times in total, total mentions reach 90 and your share of voice is 30 percent. Run the same calculation for category prompts only, and you see your potential to win new customers more clearly.
You can also weight by position: 3 points for first place, 2 for second and 1 for the rest. However, position data is so volatile that I use the weighted score only as a supporting signal. Likewise, the Citation Share metric that Bing Webmaster Tools put into preview in June 2026 follows a similar logic: it shows your site's share of all citations shown for the same query. If you struggle to pick competitors, the method in my SEO competitor analysis guide works here too.
How do you score sentiment and accuracy?
A simple three level scale works for sentiment: positive, neutral, negative. Positive means the answer recommends you or highlights a strength. Negative means it contains a warning, a complaint or labels such as “expensive” or “slow”. When in doubt, mark it neutral, because consistency matters more than precision.
For accuracy, start with a brand fact sheet. List your services, address, service area, founding year, pricing policy and current product names. Then mark each claim in an answer as correct, incomplete or wrong against that sheet.
Wrong information is the most valuable finding, because you can fix it. A wrong address, a closed branch or an old price usually comes from an outdated source somewhere online. Once you find and update that source, you can then see the difference in the next round. So always log the source URLs of wrong answers separately. Accuracy is also the part of AI visibility you can improve fastest, since the root cause often sits on a page you control.
Does ChatGPT know your brand? How do you run a 10 minute manual test?
This test does not replace the protocol, but it shows where you stand today in ten minutes. Set a timer and follow these steps:
- Setup (minute 1): Open a new Temporary Chat in ChatGPT and choose the Unpersonalized option.
- Recognition (minutes 2 and 3): Ask “What is [Brand], what does it do and where does it operate?” and compare the answer with your fact sheet.
- Source check (minute 4): In the same chat, ask it to search the web and show its sources; check whether your site is among them.
- Category question (minutes 5 and 6): In a new temporary chat, ask “Which [service] providers in [city] would you recommend?” without naming your brand.
- Comparison (minute 7): Ask “What are the differences between [Brand] and [Competitor]?” and note which side the tone favours.
- Repeat (minutes 8 and 9): Ask the category question again in two more temporary chats and count how often you appear.
- Log (minute 10): Enter the results in your tracking sheet, one row per answer.
The result usually falls into one of three scenarios. If ChatGPT knows you and names you in the category question, your foundation is solid. If it knows you but leaves you out of the category answer, then the problem is most likely authority. And if it does not know you at all, first check that your site is technically readable with our AI visibility checker. The tool scans AI bot rules in robots.txt, your llms.txt file, structured data and citable content structure, then gives you a score out of 100.
What should a monthly AI visibility report include?
The monthly report should fit on one page that both executives and specialists can read. A few clear numbers with commentary keep AI visibility on the leadership agenda far better than a pile of tables. I build the report from these sections:
- Summary: this month's value for all six metrics, the change from last month and the margin of error.
- Platform view: mention rate, citation rate and share of voice for each platform.
- Prompt type view: separate results for brand, category, comparison and problem prompts.
- Most cited pages: your own URLs and the third party sources that feed your competitors.
- Wrong information: each error, its likely source and the fix.
- Official data: Search Console impressions, Bing citations and AI referral visits.
- Next month: two or three clear actions, each with an owner.
To separate AI referral visits in analytics, use my guide on detecting AI traffic in GA4. If your team is not used to reading reports yet, how to read a digital marketing report is a good place to start. As a result, your AI visibility numbers sit at the same table as revenue and demand data.
What AI data do Search Console and Bing Webmaster Tools give you?
Google announced the generative AI performance reports in Search Console on June 3, 2026 and rolled them out to all sites as of August 31, 2026. According to the Search Console help page, the Search report shows your impressions in AI Overviews and AI Mode by page, country, device and date. However, it shows no clicks and no queries, only impressions. These impressions also still count in the overall Performance report.
One warning: sites that opt out through the generative AI control in Search Console receive no impressions from these features. So before you start, confirm that the control under Settings still sits at the default, which is inclusion. If you are new to the tool, my Search Console guide covers the basics.
On the Microsoft side, Bing Webmaster Tools launched its AI Performance report on February 10, 2026. It shows citations in Copilot and AI-generated summaries in Bing, including total citations, average cited pages, page level citation activity and a sample of grounding queries. As a result, you see roughly which of your pages serve as sources for which query families. Your own prompt set then covers what both reports miss: ChatGPT, Gemini and Perplexity.
Which tools can measure AI visibility at scale?
Measuring by hand works for small sets; as prompts and platforms grow, you need tools that measure AI visibility on a schedule. I group them into four categories:
- Official platform reports: the Search Console generative AI report and Bing Webmaster Tools AI Performance. They are free, but each covers only its own ecosystem.
- AI modules in SEO suites: Semrush AI Visibility Toolkit and Ahrefs Brand Radar fall into this group. Both sit in the same dashboard as your existing SEO data.
- Dedicated trackers: tools such as Profound, Peec AI and Otterly.AI run your prompt set across platforms on a schedule and collect mention and citation data.
- Your own setup: a spreadsheet, manual runs and, if needed, automation through APIs. It is the cheapest route, but it takes discipline.
Ask two questions before you pick a tool. How many times does it run each prompt, and does it report rates? Does it collect answers through an API or from the interface real users see? In practice, those two answers are not always the same. A tool that reports a precise rank from a single run paints a misleading picture, given the research above.
Technical readiness, however, is a separate topic. If you want to know whether AI crawlers can reach your site at all, see my guide to AI crawlers and robots.txt.
What are the most common measurement mistakes?
In my own rounds and in reports I review, these mistakes come up most often:
- Deciding on the basis of one screenshot, even though one answer is a sample, not a result.
- Changing the prompt set every month and still comparing the results.
- Measuring brand prompts only and skipping category prompts.
- Testing with personalization on, because the answer then reflects your own history.
- Making position the headline metric.
- Presenting small changes as a win or a crisis without checking the margin of error.
- Widening the competitor set mid month and misreading the drop in share of voice.
What these mistakes share is that they make measurement impossible to repeat. So the most important rule of the protocol is to ask the same question under the same conditions every month. Another common mistake is to measure AI visibility once and then drop it. AI visibility is a trend line you read over months, not a one off snapshot, so collect three months of data before you make big decisions.
How do you turn the numbers into action?
Measurement only creates value once it turns into action. First, a low mention rate usually points to weak authority and too few third party mentions. A low citation rate, on the other hand, suggests that your content is hard to quote or closed to crawlers. For wrong information, find the source and fix it first.
You will find the steps to grow visibility in my ChatGPT and Gemini visibility guide and the wider framework in what is GEO. If you have no time to run the rounds yourself, my team and I set up the prompt set and prepare the monthly report as part of our AI SEO and GEO services. Whichever route you choose, the first step is the same: run the 10 minute test this week and start the log you will use to measure AI visibility from now on.




