Skip to content
Join waitlist

How to measure AI visibility without chasing a score

Track mentions, citations, and accurate recommendations with a repeatable question set. A worked example shows why a rising score can hide an unchanged result.

An amber glass ruler rests across ivory observation cards with three amber markers on a pale stone slab.
Editorial illustration created with AI.
In this guide

To measure AI visibility, repeat a fixed set of buyer questions and record whether your brand is mentioned, linked, or recommended accurately. Track identifiable visits and useful actions separately. Keep the underlying answers so you can understand what changed before deciding what to fix.

Our view is that a visibility report should help you choose your next piece of work. A rising score can hide an inaccurate recommendation, a shift toward easier questions, or more citations that have little connection to a buying decision. The method below makes those differences visible.

Decide what you want an answer to do

Start with a business question you can act on. For a small software company, that might be whether people comparing products can discover an accurate explanation of its strongest use case.

“Increase AI visibility” leaves too much room for interpretation. “Find out whether our product is suggested for the workflows it actually supports” gives you something concrete to inspect.

Use these definitions consistently in your own reporting. They are a proposed measurement convention, rather than a universal standard across tools:

Outcome What to record What it tells you
Brand mention The answer names your business Your brand appeared in that response
Citation to your site The answer includes a clickable source link to a page you own Your page was linked as a source
Recommendation The answer explicitly suggests your product for the stated need Your product was presented as an option
Accurate recommendation The suggestion fits the need and its material product claims match verified facts The answer describes a defensible reason to consider you

An independent review can recommend your product without linking to your website. Record that source separately from citations to pages you own. An educational article can receive a citation without the answer recommending your business at all.

These outcomes can overlap. Don't add their percentages together, and don't turn a passing mention into an endorsement. Save the relevant wording whenever a classification needs judgment.

Our guide to getting your brand recommended by AI explains how to connect buying questions to evidence. Measurement starts by defining which of those questions you're trying to answer.

Build a question set you can keep using

Choose questions from available customer conversations, support requests, or sales notes. Remove private details. If you begin with questions your team proposes, label them as assumptions until you can compare them with actual customer language.

Separate learning questions from choosing questions. “How does shared scheduling work?” asks for an explanation. “Which scheduling tool supports three mobile mechanics and collects a service address?” asks for suitable options under specific constraints.

That distinction affects the denominator. Count recommendations against the questions where a recommendation is relevant. Mixing dozens of definition questions into that rate makes it harder to see how well you cover buying decisions.

For a manageable first sample, try twelve questions on one surface: six about understanding the problem and six about choosing a solution. Run each three times over a defined review period. That produces 36 planned observations. The numbers are a suggested workload, not a benchmark or a statistically representative sample.

Keep branded questions in a separate group. Asking “What does our company do?” checks brand accuracy. Asking for suitable products without naming yours checks a different discovery situation.

Save enough context to repeat the observation

For each run, keep the exact prompt, date, product or search surface, visible model label, language, relevant location, and whether search was used. Record settings you can observe, including account state and whether the conversation started fresh. Mark unavailable details as unknown.

Save the complete response and its source links, then add your classifications. A percentage without the answer behind it gives the next reviewer little to inspect.

Keep surfaces separate. Google says AI Mode and AI Overviews can use different models and techniques, so their responses and links can vary. That supports comparing each surface on its own terms. See Google's AI features documentation.

Use the same collection method for later comparisons. A change in product, prompt wording, or search setting deserves a note and may require a new baseline.

Log technical failures separately. A failed request is not evidence that your brand was absent. An evaluable answer that names other products and omits yours does count as an absence. Show both the planned and completed totals so missing runs don't quietly change the story.

Repetition helps you inspect variation, but three runs of the same prompt are still closely related observations. This sample describes the questions and conditions you tested; it doesn't establish your share of every answer real customers receive.

Read a small sample with the right denominators

Consider a fictional booking app for mobile bicycle mechanics. Assume it supports separate calendars and address collection, but has no parts inventory feature. These capabilities and all figures below are hypothetical, not HoneyWeRank results.

The team completes all 36 observations: eighteen from learning questions and eighteen from choosing questions. It records:

Measure Calculation Result within this sample
Answers mentioning the brand 18 of 36 evaluable answers 50%
Answers linking to the company's site 12 of 36 evaluable answers 33.3%
Choosing answers recommending the app 8 of 18 choosing answers 44.4%
Choosing answers with an accurate recommendation 6 of 18 choosing answers 33.3%

Two of the eight recommendations incorrectly claim the app manages parts inventory. That makes two of eight recommendations materially inaccurate, or 25% of the recommendations in this sample. The error rate uses a different denominator from the recommendation rate.

The useful next task is to inspect those two answers and their sources. If the company's own pages imply inventory support, correct them. If a third-party page is outdated, record where the claim appears and decide how to address it. If no supporting source is visible, retain the uncertainty.

A report showing only eight recommendations would miss that work. The more useful report includes the two inaccurate claims, the affected questions, and the page someone can improve.

Check whether the questions changed before celebrating growth

Now suppose the team adds six learning questions in a later period, each run three times. For this illustration, every result on the original twelve questions stays exactly the same. The eighteen new observations include fifteen citations to the company's educational pages.

The combined citation rate rises from 12 of 36 answers, or 33.3%, to 27 of 54 answers, or 50%. Accurate recommendations on the original choosing questions remain six of eighteen, or 33.3%.

The higher citation rate reflects a different question mix. It doesn't show improved performance on the original questions.

Preserve the original set as a continuing comparison group and report new questions separately. Add questions when they represent a real customer need, then establish their own baseline. Expanding coverage is useful; silently changing what a trend measures makes the trend harder to interpret.

Apply the same check when evaluating a monitoring tool. Ask which prompts entered or left the sample, how each outcome is defined, and whether you can inspect the saved answers. A score becomes useful when its movement can be traced to specific observations.

Keep visits and business actions in a separate view

An answer can mention your company without sending a visitor. A linked source can receive no clicks. Website analytics describe recorded activity on your site, so they cannot serve as a complete census of exposure inside answers.

In Google Analytics, review the Traffic acquisition report using session source or session source and medium. Inspect identifiable referring sources and the actions those sessions complete. Google's Traffic acquisition documentation explains these dimensions and the session key event rate.

Choose an action that matters to your business, such as a qualified inquiry or completed trial setup. Verify that it is recorded correctly. Keep the qualification rule stable, and show counts alongside rates when volumes are small.

Treat identifiable referrals as observed traffic, rather than all traffic influenced by AI. Someone may see a recommendation and visit later through a different route. Don't assign an unexplained rise in direct visits to AI merely because it happened during the same period.

Google also reports appearances in AI Overviews and AI Mode within Search Console's overall Web search performance. A change in total Web clicks therefore isn't a standalone measure of AI visibility. See its performance measurement guidance.

Search Console clicks and Analytics sessions are calculated differently. Compare their trends with matching filters where practical, while allowing for measurement differences. Google's guide to using Search Console and Analytics together explains why the totals need not match.

Turn the report into one specific improvement

At each review, choose a finding with evidence behind it. A missing citation, an inaccurate capability, and a poor fit for a buying question call for different responses.

If an explanation is buried, use the answer-first content guide to make it easier to find. If the page makes a claim without proof, review what makes a page worth citing. If the answer recommends a capability you don't offer, start with the incorrect statement and its source.

Write down the affected question, saved response, proposed change, responsible person, and next review point. Preserve the original page before editing it. Later, repeat the same questions and inspect whether the description changed.

A before-and-after difference can guide further investigation, but it doesn't by itself prove that your edit caused the change. Keep other changes and collection differences visible.

For your next review, bring one saved answer that deserves attention and one page your team can improve. A report that leads to a clearer product explanation has done useful work, even before the headline percentage moves.

References

About Uriel Bitton

HoneyWeRank. Prepared with AI assistance using the sources linked in the article; the worked example is hypothetical.

About our journal
Back to the journal