Growth by Design Blog | Marketing Strategy & Business Growth

How to Measure AI Visibility: The Full Method

Written by Izzy Gregorio | Sep 10, 2026, 12:30:00 PM

Pillar · Measurement

How to Measure AI Visibility

Every vendor will hand you a score and imply it is objective. None of them agree. That is not a reason to skip measurement. It is a reason to own your method.

By Izzy Gregorio  ·  Updated August 2026  ·  14 min read

 

In short

AI visibility is measured with a fixed prompt set, run repeatedly across multiple engines on a monthly cadence, scored on five metrics: share of model, citation rate, prompt coverage, competitive citation gap, and AI referral traffic. No industry-standard formula exists, so consistency of method matters more than the absolute number any single tool reports.

Start with the distinction

 

What you are actually counting

Two outcomes get conflated constantly, and the fix for each is different.

A mention is the model naming your brand inside an answer. A citation is the model naming you and linking to a source.

Mention without citation gives you influence and no traffic. Citation without mention gives you a link buried in a source list nobody reads. Semrush named the space between them in 2025 and called it the mention-source divide, finding fewer than one in five brands achieve both consistently.

A mention problem is an entity and priors problem. A citation problem is a retrieval and structure problem.

Diagnose which one you have before buying anything, because a tool that reports a single blended score will not tell you.

 

The metrics

 

Five numbers, and what each one is for

Metric What it measures What moves it
Share of model Percentage of category responses that mention your brand Entity consistency, third-party corroboration, model priors
Citation rate Of those mentions, how many link to you Extractable structure, freshness, crawler access
Prompt coverage How many relevant prompts surface you at all Cluster completeness across the five prompt types
Competitive citation gap Your share of model minus your nearest competitor's Everything above, relative to someone specific
AI referral traffic Sessions from AI platforms in analytics Lagging indicator. Undercounts badly

The fourth one is the number that gets a budget approved. The first three describe your position. The fourth names who is beating you.

A named competitor in an answer where your brand should be does more inside a leadership meeting than any percentage.

On the fifth: treat AI referral traffic as a floor, never a measure. A buyer who reads about you in an answer and later types your name arrives as direct or branded search. The referral report will never see them.

The instrument

 

Build a prompt set that means something

A score is only as good as the questions behind it. Most vendor scores are built on prompts you never chose.

Use the five prompt types from the Semrush category work. Kevin Indig's analysis applied this structure across 1,094 US categories between January and June 2026, covering more than 50,000 brands and 600,000 citations, which makes it the most defensible public structure available.

  1. 1

    Definition. What is [category]?

  2. 2

    Comparison. [Your brand] versus [named competitor]

  3. 3

    Alternatives. Alternatives to [named competitor]

  4. 4

    Use case. Best [category] for [specific customer situation]

  5. 5

    Buying. Which [category] should I buy?

Then expand. Five prompt types times four to six real customer situations produces 20 to 30 prompts, enough to be stable month over month and small enough that a person can actually run it.

Four rules that decide whether the number is trustworthy

Run each prompt three times per engine, not once. AirOps found only 30 percent of brands remain visible across consecutive AI responses. Ask the same question twice and the answer names different brands roughly seven times out of ten. A single run is noise. Record the rate at which you appear, not whether you appeared.

Run logged out. Personalization will flatter you and the flattery is not reproducible.

Freeze the wording. Changing a prompt changes the result. A prompt set you edit monthly measures your editing, not your visibility.

Record the date and the engine. Every number is a snapshot of a system that has already moved.

The first rule is the one most methods skip, including several commercial ones. A score built on single runs of each prompt is measuring volatility as much as visibility, which is why two tools can report different numbers for the same brand on the same day and both be honest.

 

The cadence

 

Monthly, not quarterly

13 weeks

Half of all AI-cited content is younger than this

3x

More likely to lose citations entirely if a page goes three months without an update

Amsive, 2026, and AirOps State of AI Search, 2026

Do the arithmetic on that against a quarterly reporting rhythm. By the time a quarterly report reaches the room, half the content currently being cited in your category did not exist when the data was collected. You are presenting history and calling it status.

Traditional organic reporting assumes a much slower clock, which is why quarterly worked there and does not work here.

Monthly measurement on a frozen prompt set is the correct rhythm, and the second month costs a fraction of the first once the instrument exists.

Why nothing agrees

 

Two tools, two numbers, same brand

This confuses people into thinking one vendor is lying. Usually neither is.

There is no standard formula. The share-of-voice calculation in circulation is straightforward, brand citations divided by total category citations times 100, but every input to it is a judgment call. Which prompts. How many. Which engines. How many times each prompt is run. Whether a mention without a link counts. Whether repeated mentions in one answer count once or twice. Whether the answer was generated logged in or logged out.

Semrush, Conductor, HubSpot's grader, and Profound each answer those questions differently, and none of them are wrong. As of August 2026, no independent cross-platform accuracy benchmark had been published for any of them. That may change, and this page will be updated when it does.

A published figure you cannot defend is worse than no figure at all.

So publish the method alongside the number. Which prompts, which engines, how many runs, which date, run logged out. A score with a documented method survives being questioned. A score without one does not.

 

Translate it

 

What to report, by who is reading it

If you are reporting to What they actually want to know The number to lead with
A board or investor group Are we losing ground to a named competitor Competitive citation gap, with the competitor named
A CEO or founder Is this real or is it a trend Share of model trend over three months, plus one verbatim answer showing a competitor named instead
A CFO reviewing spend What does this cost and what returns Prompt coverage against pipeline-relevant queries, plus AI referral traffic as a floor
A marketing team executing What do we fix on Monday Citation rate by page, which points directly at structure
An agency partner Are they doing the work The prompt set itself, frozen, with dated month-over-month results

If an agency cannot show you the prompt set, they are not measuring. They are reselling a dashboard.

The framework

 

The four-layer measurement stack

Ordered deliberately. Each layer is worthless without the one below it.

  1. 1

    Access

    Can the engines read you at all? Crawler permissions, indexation, rendering. Binary, and free to check.

  2. 2

    Presence

    Do you appear? Share of model and prompt coverage on a frozen set, run repeatedly.

  3. 3

    Attribution

    When you appear, do you get credit? Citation rate, and whether the link resolves.

  4. 4

    Position

    Are you ahead or behind? Competitive citation gap against named rivals.

Most brands attempt layer four first, because it is the interesting one. Then they discover the number is unstable, because layers one through three were never established and the score is measuring noise.

One question worth sitting with. Your last marketing report contained a number that somebody in the room could have challenged. If they had asked how it was calculated, which prompts produced it, how many times each was run, and on what date, would the answer have been available in the room? That is the whole standard, and almost nothing in this category currently meets it.

 

Common questions

 

Measuring AI visibility, answered

What is share of model?

Share of model is the percentage of category responses across AI engines that mention your brand. It measures presence rather than traffic. It is calculated by running a fixed prompt set, counting how many responses name you, and dividing by total responses. Method transparency matters more than the number.

What is a good AI citation rate?

No credible benchmark exists yet, because no independent cross-platform accuracy standard has been published. Semrush's 2025 work found fewer than one in five brands achieve both frequent mentions and consistent citations, which suggests most brands sit low. Track your own trend rather than chasing an industry figure.

How many prompts do I need to measure reliably?

Twenty to thirty is a workable floor for a single category, from five prompt types applied across four to six real customer situations. Run each one at least three times per engine rather than once, since AirOps found only 30 percent of brands remain visible across consecutive AI responses. A single run measures volatility as much as visibility.

Why do two AI visibility tools give me different scores?

Because each one chooses different prompts, engines, run counts, aggregation rules, and mention definitions, and no industry-standard formula exists. Answer volatility compounds it: only 30 percent of brands remain visible across consecutive responses, so two tools can query the same engine on the same day and honestly get different results. Pick one method, document it, and apply it identically every month.

How often should I measure AI visibility?

Monthly. Half of all AI-cited content is less than 13 weeks old per Amsive's 2026 analysis, and AirOps found pages that go three months without an update are more than three times as likely to lose citations entirely. Quarterly reporting describes a position that has substantially turned over since collection. Monthly measurement on a frozen prompt set is also inexpensive after the first run.

Is AI referral traffic a good measure?

It is a floor, not a measure. It undercounts badly, because a buyer who encounters you in an AI answer and later searches your name arrives as direct or branded traffic. Use it to confirm direction, never as the headline number in a report.

Do I need a paid platform to measure this?

Not to start. A frozen prompt set, four engines, a logged-out browser, and a spreadsheet produces a defensible baseline. Budget more than thirty minutes, since running each prompt three times per engine is what makes the result stable. Buy tooling once you know the shape of the problem and need the monthly run automated.

Get the data

 

The Citation Gap Report 2026

The full methodology, the prompt set structure, and the citation data behind everything on this page. Numbers you can check, sources you can trace, and a measurement method you can apply to your own category this week.

Download the Citation Gap Report