How to Measure AI Visibility: The Full Method
Pillar · Measurement
How to Measure AI Visibility
Every vendor will hand you a score and imply it is objective. None of them agree. That is not a reason to skip measurement. It is a reason to own your method.
By Izzy Gregorio · Updated August 2026 · 14 min read
In short
AI visibility is measured with a fixed prompt set, run repeatedly across multiple engines on a monthly cadence, scored on five metrics: share of model, citation rate, prompt coverage, competitive citation gap, and AI referral traffic. No industry-standard formula exists, so consistency of method matters more than the absolute number any single tool reports.
Start with the distinction
What you are actually counting
Two outcomes get conflated constantly, and the fix for each is different.
A mention is the model naming your brand inside an answer. A citation is the model naming you and linking to a source.
Mention without citation gives you influence and no traffic. Citation without mention gives you a link buried in a source list nobody reads. Semrush named the space between them in 2025 and called it the mention-source divide, finding fewer than one in five brands achieve both consistently.
A mention problem is an entity and priors problem. A citation problem is a retrieval and structure problem.
Diagnose which one you have before buying anything, because a tool that reports a single blended score will not tell you.
The metrics
Five numbers, and what each one is for
| Metric | What it measures | What moves it |
|---|---|---|
| Share of model | Percentage of category responses that mention your brand | Entity consistency, third-party corroboration, model priors |
| Citation rate | Of those mentions, how many link to you | Extractable structure, freshness, crawler access |
| Prompt coverage | How many relevant prompts surface you at all | Cluster completeness across the five prompt types |
| Competitive citation gap | Your share of model minus your nearest competitor's | Everything above, relative to someone specific |
| AI referral traffic | Sessions from AI platforms in analytics | Lagging indicator. Undercounts badly |
The fourth one is the number that gets a budget approved. The first three describe your position. The fourth names who is beating you.
A named competitor in an answer where your brand should be does more inside a leadership meeting than any percentage.
On the fifth: treat AI referral traffic as a floor, never a measure. A buyer who reads about you in an answer and later types your name arrives as direct or branded search. The referral report will never see them.
The instrument
Build a prompt set that means something
A score is only as good as the questions behind it. Most vendor scores are built on prompts you never chose.
Use the five prompt types from the Semrush category work. Kevin Indig's analysis applied this structure across 1,094 US categories between January and June 2026, covering more than 50,000 brands and 600,000 citations, which makes it the most defensible public structure available.
- 1
Definition. What is [category]?
- 2
Comparison. [Your brand] versus [named competitor]
- 3
Alternatives. Alternatives to [named competitor]
- 4
Use case. Best [category] for [specific customer situation]
- 5
Buying. Which [category] should I buy?
Then expand. Five prompt types times four to six real customer situations produces 20 to 30 prompts, enough to be stable month over month and small enough that a person can actually run it.
Four rules that decide whether the number is trustworthy
Run each prompt three times per engine, not once. AirOps found only 30 percent of brands remain visible across consecutive AI responses. Ask the same question twice and the answer names different brands roughly seven times out of ten. A single run is noise. Record the rate at which you appear, not whether you appeared.
Run logged out. Personalization will flatter you and the flattery is not reproducible.
Freeze the wording. Changing a prompt changes the result. A prompt set you edit monthly measures your editing, not your visibility.
Record the date and the engine. Every number is a snapshot of a system that has already moved.
The first rule is the one most methods skip, including several commercial ones. A score built on single runs of each prompt is measuring volatility as much as visibility, which is why two tools can report different numbers for the same brand on the same day and both be honest.
The cadence
Monthly, not quarterly
Half of all AI-cited content is younger than this
More likely to lose citations entirely if a page goes three months without an update
Amsive, 2026, and AirOps State of AI Search, 2026
Do the arithmetic on that against a quarterly reporting rhythm. By the time a quarterly report reaches the room, half the content currently being cited in your category did not exist when the data was collected. You are presenting history and calling it status.
Traditional organic reporting assumes a much slower clock, which is why quarterly worked there and does not work here.
Monthly measurement on a frozen prompt set is the correct rhythm, and the second month costs a fraction of the first once the instrument exists.
Why nothing agrees
Two tools, two numbers, same brand
This confuses people into thinking one vendor is lying. Usually neither is.
There is no standard formula. The share-of-voice calculation in circulation is straightforward, brand citations divided by total category citations times 100, but every input to it is a judgment call. Which prompts. How many. Which engines. How many times each prompt is run. Whether a mention without a link counts. Whether repeated mentions in one answer count once or twice. Whether the answer was generated logged in or logged out.
Semrush, Conductor, HubSpot's grader, and Profound each answer those questions differently, and none of them are wrong. As of August 2026, no independent cross-platform accuracy benchmark had been published for any of them. That may change, and this page will be updated when it does.
A published figure you cannot defend is worse than no figure at all.
So publish the method alongside the number. Which prompts, which engines, how many runs, which date, run logged out. A score with a documented method survives being questioned. A score without one does not.
Translate it
What to report, by who is reading it
| If you are reporting to | What they actually want to know | The number to lead with |
|---|---|---|
| A board or investor group | Are we losing ground to a named competitor | Competitive citation gap, with the competitor named |
| A CEO or founder | Is this real or is it a trend | Share of model trend over three months, plus one verbatim answer showing a competitor named instead |
| A CFO reviewing spend | What does this cost and what returns | Prompt coverage against pipeline-relevant queries, plus AI referral traffic as a floor |
| A marketing team executing | What do we fix on Monday | Citation rate by page, which points directly at structure |
| An agency partner | Are they doing the work | The prompt set itself, frozen, with dated month-over-month results |
If an agency cannot show you the prompt set, they are not measuring. They are reselling a dashboard.
The framework
The four-layer measurement stack
Ordered deliberately. Each layer is worthless without the one below it.
- 1
Access
Can the engines read you at all? Crawler permissions, indexation, rendering. Binary, and free to check.
- 2
Presence
Do you appear? Share of model and prompt coverage on a frozen set, run repeatedly.
- 3
Attribution
When you appear, do you get credit? Citation rate, and whether the link resolves.
- 4
Position
Are you ahead or behind? Competitive citation gap against named rivals.
Most brands attempt layer four first, because it is the interesting one. Then they discover the number is unstable, because layers one through three were never established and the score is measuring noise.
One question worth sitting with. Your last marketing report contained a number that somebody in the room could have challenged. If they had asked how it was calculated, which prompts produced it, how many times each was run, and on what date, would the answer have been available in the room? That is the whole standard, and almost nothing in this category currently meets it.
Common questions
Measuring AI visibility, answered
What is share of model?
Share of model is the percentage of category responses across AI engines that mention your brand. It measures presence rather than traffic. It is calculated by running a fixed prompt set, counting how many responses name you, and dividing by total responses. Method transparency matters more than the number.
What is a good AI citation rate?
No credible benchmark exists yet, because no independent cross-platform accuracy standard has been published. Semrush's 2025 work found fewer than one in five brands achieve both frequent mentions and consistent citations, which suggests most brands sit low. Track your own trend rather than chasing an industry figure.
How many prompts do I need to measure reliably?
Twenty to thirty is a workable floor for a single category, from five prompt types applied across four to six real customer situations. Run each one at least three times per engine rather than once, since AirOps found only 30 percent of brands remain visible across consecutive AI responses. A single run measures volatility as much as visibility.
Why do two AI visibility tools give me different scores?
Because each one chooses different prompts, engines, run counts, aggregation rules, and mention definitions, and no industry-standard formula exists. Answer volatility compounds it: only 30 percent of brands remain visible across consecutive responses, so two tools can query the same engine on the same day and honestly get different results. Pick one method, document it, and apply it identically every month.
How often should I measure AI visibility?
Monthly. Half of all AI-cited content is less than 13 weeks old per Amsive's 2026 analysis, and AirOps found pages that go three months without an update are more than three times as likely to lose citations entirely. Quarterly reporting describes a position that has substantially turned over since collection. Monthly measurement on a frozen prompt set is also inexpensive after the first run.
Is AI referral traffic a good measure?
It is a floor, not a measure. It undercounts badly, because a buyer who encounters you in an AI answer and later searches your name arrives as direct or branded traffic. Use it to confirm direction, never as the headline number in a report.
Do I need a paid platform to measure this?
Not to start. A frozen prompt set, four engines, a logged-out browser, and a spreadsheet produces a defensible baseline. Budget more than thirty minutes, since running each prompt three times per engine is what makes the result stable. Buy tooling once you know the shape of the problem and need the monthly run automated.
Get the data
The Citation Gap Report 2026
The full methodology, the prompt set structure, and the citation data behind everything on this page. Numbers you can check, sources you can trace, and a measurement method you can apply to your own category this week.