CLIENT REVENUE GENERATED$9,720,200+Independent consulting · Since 2008
Talk Strategy

Analytics & Conversions / Practical guide

How to Measure AI Search Visibility: Mentions, Citations and Leads

Measure AI mentions, citations and recommendations separately. Use an observation log, clear denominators and careful referral and lead reporting.

1. Decide which outcome you are measuring

When someone says their business is visible in AI search, I ask what they saw. Was the company named? Was its website cited? Was the recommendation accurate? Did anyone visit or enquire? Each observation can be useful, but combining them into one score can hide the work the business actually needs.

OutcomeEvidenceWhat it does not prove
MentionThe business is named in an answer.A link, endorsement or visit
CitationA visible source points to a specific URL.That a user clicked it
RecommendationThe answer explicitly presents the business as an option for the task.A verified endorsement, visit or qualified enquiry
ReferralA visit is recorded with an identifiable source.A qualified buyer
Qualified enquiry / business outcomeAn enquiry or sale can be connected to the visit.The complete influence of every earlier interaction
Five outcomes I keep separate

I also record inaccurate descriptions. Being mentioned for a service you do not offer can create the wrong enquiries. An AI visibility audit should examine the quality of the representation, not just whether the name appears.

An answer might name your business without linking to it, or link to a useful page without recommending you. I would keep those observations separate. Agree what counts as a recommendation, save the wording and check whether the description is accurate.

Analyst workspace with comparison notes and an abstract conversation on a laptop.

2. Build a question set around the buying decision

I start with the services, products and markets the business can genuinely serve. Then I collect the questions buyers ask while discovering options, comparing suppliers and checking fit. A branded question belongs in its own group because it is a different test from a question that never names the business.

For a specialist software provider, I might test the problem the software solves, the integrations a buyer needs and the implementation concerns raised before a demonstration. I would not fill the list with dozens of slight rewrites of “best software company” just to create a larger report.

I keep a fixed core set for comparison and a separate exploratory list for new questions. Otherwise, changing the questions every month can make a result look like progress when the test itself has changed. That list tells me what we tested. It does not tell me how many people asked those questions.

3. Make the observation repeatable

For each test, I record the exact question, platform, date, language, market setting where available and whether a fresh conversation was used. I retain the visible answer and source URLs. If a follow-up question was involved, I keep that context too.

  • Keep the conditions clear: compare the same question groups under documented conditions.
  • Inspect the source: check whether the citation points to the intended page and supports the statement.
  • Repeat observations: distinguish an isolated appearance from a recurring pattern.
  • Record absence: a test without a mention belongs in the denominator.
  • Separate platforms: do not treat different answer systems as one identical ranking list.

Here is an illustrative calculation: if a brand appears in 6 of 20 recorded answers, that is a 30% mention rate within that sample. It is not 30% of all AI searches, and it says nothing about clicks. Without the dates and question group beside it, I would not use that number to claim progress.

01Observe

Mention or linked citation?

02Connect

Was a visit recorded?

03Evaluate

Did it lead to useful action?

4. Use a log that someone else can interpret

A useful log preserves the question and conditions behind the result. The fields below form a copyable template: create one row per question, product and observation time. Keep screenshots and full responses in a private evidence folder rather than putting customer information into a public report.

FieldWhat to record
Run ID and question groupA stable identifier; branded and non-branded questions remain separate.
Exact prompt and contextThe verbatim question, any follow-ups and whether the conversation was fresh.
Product, mode and test surfaceConsumer interface or API; product/mode/model where shown; search enabled where applicable.
Timestamp and localeDate, time zone, language and market setting where available. Mark unavailable settings as unknown.
Answer statusUsable answer, no AI answer shown, unavailable, timed out or incomplete. Do not turn missing evidence into a negative brand result.
Brand mentionYes/no for an inspected answer; unknown if it could not be inspected. Retain the wording.
Linked citation and URLWhether a visible citation points to the monitored domain, plus every relevant cited URL.
RecommendationWhether the answer explicitly presents the business as an option for the task. A name in a list is not automatically an endorsement.
Source support and business accuracySupported, partly supported, unsupported or not assessed; explain whether the source supports the relevant statement.
Evidence and change notesA private screenshot/response reference, collection limitations and any changed test condition.
Observation log: fields to copy into your working sheet

A filled row, without invented performance claims

The following example is synthetic, not a real platform answer or client record. Example Provider is a fictional business; the example URL is not evidence of a real service.

FieldIllustrative entry
Question and group“Which providers offer migration support for a WooCommerce store?” / non-branded service selection
Product and conditionsIllustrative search-enabled assistant, consumer UI, fresh conversation, English, India; no actual product tested
Observation date5 October 2026, 10:00 IST; illustrative timestamp
Answer statusUsable answer in this fictional example
Mention / citation / recommendationMention: yes. Citation: yes, https://example.com/migration/. Recommendation: no; the fictional wording describes a service but does not recommend the provider.
Source and notesSource support: not assessed. No screenshot exists because this is a demonstration of the log, not a completed test.
Synthetic observation: EX-001

The important distinction is that the three observation fields can differ. A cited page may provide background information without the business being recommended. Inspect the wording before assigning a result.

5. Show the denominator beside every percentage

Synthetic worked example: suppose you schedule 20 checks in one documented product and mode. Sixteen return usable answers; two time out and two are unavailable. Four of the 16 usable answers cite your domain. Six mention the business and two explicitly recommend it. These events may overlap; they are not separate buckets to add together.

MeasureCalculationWhat it describes
Observation coverage16 usable / 20 scheduled = 80%How much of the scheduled test produced inspectable evidence
Citation incidence4 citing answers / 16 usable answers = 25%Citation frequency within the inspected sample
Mention incidence6 mentioning answers / 16 usable answers = 37.5%Name appearances within the same inspected sample
Recommendation incidence2 recommending answers / 16 usable answers = 12.5%Explicit recommendations within that sample
Illustrative calculations, not SEOFreelance.net results

Report the four missing checks and why they are missing. An inspected answer that does not mention you is a valid negative observation and stays in the denominator. An unavailable answer is not the same thing. None of these percentages measures market share, search demand or website clicks.

AI Overview appearance needs a separate measure

For a separate illustrative Google Search test, imagine 20 completed searches, five AI Overviews appearing, and two of those five Overviews citing your domain. Overview appearance is 5/20 = 25%. Citation incidence among appearing Overviews is 2/5 = 40%. You may also report 2/20 = 10% of completed searches produced an Overview citation, but name that measure explicitly.

A completed search with no AI Overview is not a failed request. Keep that status distinct from an error or an uncompleted check. Google explains that AI Overviews do not appear for every query and that their sources can differ from AI Mode. See Google’s explanation of its AI features.

Do not mix test surfaces or quietly replace questions

An API run is a separate test surface from a consumer interface. Record the endpoint/model settings and retrieval tools for API tests rather than treating their answers as consumer-product observations. If a tool, mode, market or question set changes, annotate the change and avoid presenting the two samples as an unchanged comparison.

For the SF AI Score framework or any other aggregate score, retain the version, component definitions and underlying observations. A summary can aid a discussion, but it cannot recover missing evidence or establish why a citation changed.

6. Do not use crawler activity as citation evidence

A request log tells you that a request reached a particular layer and received a response. It does not show that the content was cited in a visible answer. A User-Agent label alone does not authenticate the requester, and even a genuine crawler fetch does not establish a recommendation, click or enquiry.

Keep crawler access monitoring separate from answer observations. If you can see fetches but no citations, inspect the actual questions and sources before calling that a content failure. Likewise, a traffic decline in a bot dashboard is not automatically a decline in buyer visibility.

7. Connect observations with analytics carefully

I review identifiable AI referral traffic in analytics, then examine landing pages, engagement and agreed conversion events. Some journeys will not carry a clear source. I do not relabel unexplained direct traffic as AI traffic just because visibility observations improved during the same period.

Google’s AI features guidance explains that AI Overviews and AI Mode appearances are included in Search Console’s overall Web performance reporting. That is not a separate, complete AI attribution report. I keep that distinction when discussing organic search changes with clients.

For enquiries, a short optional “How did you hear about us?” question can add context, but I label it as self-reported information. It can complement the tracked journey rather than replace it. My analytics work focuses on a defensible connection between discovery and useful action, including the gaps we cannot resolve.

When reviewing GA4, label the scope of the traffic-source dimension you used. Session-source reporting and event-scoped attribution answer different questions. Keep the source/medium rules, landing pages and agreed lead definition beside the report, rather than treating every visit from a known assistant domain as a qualified buyer. Google documents channel groups and their scopes.

Exclude your own tests and known spam from lead totals. Store contact details only in the approved CRM or enquiry system; use aggregate counts or pseudonymous references in the observation report.

8. Turn the report into a useful next step

If the business is repeatedly described incorrectly, I review its service information and the sources contributing to that description. If an answer cites a weak or outdated page, I inspect the page’s accuracy, usefulness and links. If visits arrive but do not convert, the next task may be the offer or enquiry experience rather than more visibility tests.

Google says there is no special AI schema required for its Search AI features. I therefore prioritise accessible, accurate content and consistent business information over selling an extra markup file as a shortcut. That guidance is specific to Google’s Search features; I do not turn it into a claim about every AI product.

My AI citation work and ongoing monitoring serve different purposes. One addresses how information can be supported and discovered; the other checks what is actually observed over time. The monthly discussion should end with a decision, an owner and a way to assess the change.

Sources and limitations

Platform guidance was checked on 5 October 2026. The examples illustrate how to record and calculate visibility; they are not client results. AI answers can vary between checks, so treat each observation as a snapshot, not a permanent result.

Common questions

Can one AI visibility score tell me whether the work is succeeding?

It can summarise a defined test, but I want the question set, sample size, platform and scoring method beside it. Without those details, comparisons can be misleading.

Is a citation better than a mention?

A relevant citation provides a route to a source, but the business goal still matters. I check whether the answer is accurate, whether the source is useful and whether any meaningful visits or enquiries follow.

How often should we run the tests?

I choose a cadence that fits the decision and the size of the question set. Consistent observations around meaningful changes are more useful than repeatedly testing a handful of prompts without a reporting purpose.

Can you guarantee that a page will be cited?

I can improve the information, accessibility and measurement, then report the observed results. Selection belongs to the platform, so I do not sell a guaranteed citation as a deliverable.

Want a clearer view of AI discovery?

I can define a relevant question set, assess your current representation and connect the findings with your website priorities.

Explore an AI visibility audit