How-To · Do It Yourself
You do not need to buy a platform to find out you are invisible. Four engines, one prompt set, a spreadsheet.
By Izzy Gregorio · Updated August 2026 · 10 min read
In short
An AI visibility test requires four engines, a fixed set of 25 prompts, a logged-out browser, and a spreadsheet with five columns. Thirty minutes produces a diagnosis: whether your brand appears, whether it gets linked, and which competitors are named instead. Buy tooling after you know the shape of the problem, not before.
Before you start
A private or incognito browser window
Accounts logged out everywhere
ChatGPT, Claude, Gemini, and Perplexity open in tabs
A spreadsheet with five columns: prompt, engine, mentioned, cited with link, competitors named
Three to five real competitor names, written down before you start
That last item is the one people skip. Choosing competitors after seeing the results is how a test becomes a justification.
Step 1
Five prompt types across five real customer situations.
| Type | Template | What it reveals |
|---|---|---|
| Definition | What is [category]? | Whether you are associated with the category at all |
| Comparison | [Your brand] versus [competitor] | Whether you are described accurately |
| Alternatives | Alternatives to [competitor] | Whether you reach buyers actively shopping away from a rival |
| Use case | Best [category] for [situation] | Whether you own a specific, defensible niche |
| Buying | Which [category] should I buy? | Whether you make the shortlist, which is one to three names |
Write them in your buyer's language, not your industry's. If your customers say cheap, write cheap. Category jargon produces a clean result about a question nobody asks.
Step 2
Open a private window. Confirm no account is signed in on any of the four engines.
A signed-in session personalizes on your history, and your history is full of your own brand. The result will flatter you and nobody else will ever see it. That is worse than no test, because it produces false confidence and it is not reproducible next month.
Run every prompt in every engine. Twenty-five prompts times four engines is 100 responses, and it moves faster than it sounds once you are in rhythm.
Run all four, because they genuinely disagree. SOCi's 2026 Local Visibility Index, covering more than 350,000 locations across 2,751 brands, found ChatGPT surfaces a local business 1.2 percent of the time, Perplexity 7.4 percent, and Gemini 11 percent.
Same businesses, three very different thresholds. One engine gives you one engine's answer.
Step 3
Mentioned. Was your brand named anywhere in the answer? Yes or no.
Cited with link. Was there a link resolving to your domain? Yes or no. This is a separate column for a reason. Mention without citation gives you influence and no traffic.
Competitors named. Write every competitor name that appeared. This is the most valuable column in the sheet and the one people leave blank.
Date and engine. Every result is a snapshot of a system that has already moved.
Do not summarize as you go. Log raw, judge later. A test where you decide what counts while you are running it is a test measuring your mood.
Step 4
Share of model. Responses that mentioned you, divided by 100 total responses.
Citation rate. Responses that linked to you, divided by responses that mentioned you.
Competitive citation gap. Your share of model minus your top competitor's, calculated against your fixed competitor list only.
The third number is the one that moves a budget.
We appear in 11 percent of category answers. That is information.
Our nearest competitor appears in 34 percent and we appear in 11. That is a decision.
Step 5
The pattern points at which of three problems you actually have, and they carry three different budgets.
| Pattern | Diagnosis | First fix |
|---|---|---|
| Zero mentions anywhere, all engines | Access or entity problem | Check robots.txt and listing consistency before writing anything |
| Mentioned often, rarely linked | Structure problem. Pages are not extractable | Move answers to the top, cut backward dependencies |
| Linked when mentioned, rarely mentioned | Entity and corroboration problem | Third-party citation supply. Coverage, listings, original data |
| Present in definition prompts only | Known but not shortlisted | Commercial page structure plus use case content |
| Present in one engine only | Engine-specific gap, often Bing indexation | Submit sitemap to Bing Webmaster Tools |
The test does not tell you why. It tells you which why to investigate.
That is still worth thirty minutes, because the three fixes cost wildly different amounts and guessing between them is how budget gets wasted.
One honest limit
A single pass through 100 responses gives you a diagnosis. If you are absent across all of them, you are absent, and the table above tells you where to look first.
What it does not give you is a number stable enough to compare against next month. AirOps found only 30 percent of brands remain visible across consecutive AI responses. Ask the same question twice and roughly seven times out of ten a different set of brands comes back.
| Version | What it costs | What it gives you |
|---|---|---|
| One run, 100 responses | Thirty minutes | A diagnosis. Am I invisible, and where should I look first |
| Three runs, 300 responses | An afternoon | A measurement. A number stable enough to trend month over month |
Run the thirty-minute version first. It answers the only question that matters right now, which is whether there is a problem at all. Upgrade to three runs the month you start reporting the number to somebody who will compare it against last month's.
Translate it
| If you are the | What to bring to the next meeting | The ask |
|---|---|---|
| CMO | The competitive citation gap, with the competitor named | Budget for the specific fix the diagnosis points at |
| Marketing director | Three verbatim answers where a competitor was named instead | Agreement that the current report is incomplete |
| Founder holding marketing | One screenshot of the buying prompt result | A decision on whether this becomes a priority this quarter |
| Agency partner | The frozen prompt set | A commitment to run it monthly and report against it |
Do this part
Save the prompt list verbatim in a document nobody edits casually. Same wording, same order, same engines, every month.
A prompt set that changes between runs produces a trend line that measures your editing. The value of this test is almost entirely in repetition, and month two costs a fraction of what month one did because the hard part is already done.
Set a monthly reminder. Half of all AI-cited content is less than 13 weeks old, per Amsive's 2026 analysis, which means quarterly testing describes a position that has substantially turned over since you collected it.
One question worth sitting with. Thirty minutes gets you a defensible baseline, a named competitor, and a diagnosis pointing at one of three budgets. Nobody needs approval to run it and nobody needs to buy anything.
So when this quarter ends without it having been run, what will the actual reason have been?
Common questions
Run a fixed set of about 25 prompts across ChatGPT, Claude, Gemini, and Perplexity in a logged-out browser. Log whether your brand was mentioned, whether a link resolved to your domain, which competitors were named, and the date. That produces a defensible diagnosis in roughly thirty minutes.
A signed-in session personalizes results based on your history, which is full of your own brand. Logged-in results flatter you, are not reproducible, and are not what a prospective customer would see. Logged-out runs are the only version that can be trended month over month.
A mention is the model naming your brand in an answer. A citation is naming you and linking to a source. Mention without citation gives influence and no traffic. Log them in separate columns, because the fix for each is different and a blended score hides which one you have.
One run is enough for a diagnosis. It is not enough for a number you will compare against next month. AirOps found only 30 percent of brands remain visible across consecutive AI responses, so a single pass carries real volatility. Run each prompt three times per engine once you start reporting the figure to somebody who will trend it.
About 25 prompts across four engines produces 100 responses, which is enough to diagnose and small enough to run by hand. One engine gives one engine's answer. SOCi's 2026 index found ChatGPT surfaces a local business 1.2 percent of the time against Gemini at 11 percent, so single-engine testing produces a badly incomplete picture.
Check access first, before writing any content. Open your robots.txt and confirm AI crawlers are not disallowed, then verify your business information is consistent across listings. Total absence across all engines usually indicates an access or entity problem rather than a content problem. It is also worth knowing that Semrush found 53.7 percent of 1,094 US categories had no consistent brand leader at all, so absence is common and the position is often still available.
Eventually, if you want the monthly run automated and historical data stored without manual work. Not to start. Running it by hand once tells you the shape of the problem, which is what makes a tooling decision informed rather than speculative.
Get the benchmark
You will have a number in thirty minutes. This tells you what it means. The full citation dataset, the methodology behind the benchmarks, and what typical share of model looks like across categories.