Selection · Nine Questions
Certifications do not sort this market, because nobody has been doing it long enough to be certified in it. What sorts it is whether a provider can produce artifacts.
By Izzy Gregorio · Updated August 2026 · 12 min read
In short
Screen GEO providers on evidence rather than credentials, because 88 percent of organizations claim GEO capability and 37 percent admit their own offering is loosely defined, per GNW Consulting and Demand Metric's 2026 survey of 225 business-to-business leaders. Nine questions separate providers who run these programs from providers who resell dashboards, and all nine are answerable in one call.
State the criteria first
Selection criteria before judging, because a screening list with no stated basis is just opinion.
Each question below meets three tests. It is answerable in a single call without technical background. A provider actually doing the work can answer it immediately from an existing artifact. And a provider repackaging SEO cannot answer it without stalling.
Questions that failed those tests and were left out: years in business, client logos, team size, certifications, and case studies.
All five are easy to assemble and none of them distinguish anyone in a category this young.
The nine
Pass: an actual file, frozen, with dates. Ideally a redacted example from a live client.
Fail: an explanation of the methodology instead of the artifact, or a reason it cannot be shared.
This is the highest-yield question in the set. Measuring AI visibility requires a prompt set to exist as a file. If it does not exist, nothing consistent is being measured.
Pass: four or more named, with a reason for the weighting tied to your buyers.
Fail: ChatGPT only, or a claim that all engines behave similarly.
Engine behavior diverges sharply. SOCi's 2026 Local Visibility Index found ChatGPT surfaces a local business 1.2 percent of the time against Gemini at 11 percent, so single-engine testing produces single-engine conclusions.
Pass: a named method with published dimensions and weightings.
Fail: we use an industry-standard score.
There is no industry standard. As of August 2026 no independent cross-platform accuracy benchmark had been published, so any provider claiming a standard is either mistaken or hoping you will not ask.
Pass: a real number, typically 30 to 50 percent, covering listings, coverage, and citation supply.
Fail: near zero, or confusion about why it matters.
Roughly 82 percent of AI citations come from earned media, per Muck Rack's December 2025 analysis. A scope that is entirely on-site content is aimed at the smaller share of citations.
Pass: before any production work, delivered as a document with the prompt set, the engines, the date, and raw results.
Fail: we establish that as we go.
Without a Day 0 number on a frozen set, month six produces an argument rather than a report.
Pass: baseline test, crawler access verification, indexation check, entity and listing consistency, prioritized fix list.
Fail: content production in week one.
Producing content before diagnosis is prescribing before examining.
Pass: citation frequency, citation accuracy, AI referral traffic in analytics, and share-of-voice movement against named competitors, with re-test dates.
Fail: traffic, rankings, or increased visibility.
If the contract measures rankings, you have bought an SEO retainer with new packaging.
Pass: an immediate, specific list. Third-party publishing timelines. Specific engine behavior. Compression of citation supply below a few months.
Fail: nothing, or a deflection.
A provider who has run these programs has hit the limits. One with no limitations to name has not.
Pass: they will say so and recommend a project instead of a retainer.
Fail: every diagnosis leads to the same twelve-month engagement.
This question tests incentive alignment more than capability, and the answer tends to be revealing.
Translate it
| Type | What they actually sell | Right for | Wrong for |
|---|---|---|---|
| Monitoring reseller | A platform subscription with a monthly report layered on | Teams with in-house execution who need measurement only | Anyone without capacity to act on findings. Most likely to be mistaken for a full service |
| Repackaged SEO agency | The previous scope, renamed | Nobody, until they build the off-site capability. Some are honest about being early, and those are workable | Anyone who needs citation supply, which is most of the market |
| Specialist boutique | Named methodology, proprietary scoring, multi-engine work, real off-site scope | Mid-market brands who need execution rather than staff augmentation | Enterprises needing multi-market coverage at scale |
| Enterprise integrated | GEO folded into a large PR, content, and technical program | Organizations above $100M with multi-category needs | Anyone under roughly $20M. The overhead is not recoverable at that size |
The clearest loser here is the repackaged SEO agency claiming full capability, and it is the most common type you will encounter.
The distinction that matters is not experience level. It is whether inexperience is disclosed.
The pricing check
Published market compilations put mid-market retainers at $5,000 to $10,000 per month, boutique specialists at $5,000 to $15,000, and audits at $1,500 to $5,000. Treat those as orientation rather than benchmarks.
Below roughly $1,500 per month, the market openly describes offerings as repackaged SEO with an AI label. That is arithmetic rather than judgment: four engines, a maintained prompt set, structural editing, and any real off-site work do not fit inside the hours that price buys.
A price far below the market range is a scope statement, whether or not anyone says so.
Equally, a price far above the range should be itemized. Ask what specifically costs more, and whether it is multi-market coverage, original data publishing, or overhead you are absorbing.
Before you sign
The frozen prompt set, attached as an exhibit. It is the instrument, and it should survive the relationship.
The Day 0 baseline, delivered before production work begins.
Named re-test dates, typically Day 45 and Day 90, with a before-and-after comparison committed in writing.
The metrics, specifically citation frequency, citation accuracy, AI referral traffic, and share-of-voice movement against named competitors.
If those four are in the agreement, the engagement is measurable regardless of who delivers it. If none of them are, you have bought effort rather than outcomes.
One question worth sitting with. Take the nine questions to your existing agency before you take them to a new one.
If they answer eight of nine well, switching costs you a working relationship for a marginal gain. If they cannot answer question one, you learned that this afternoon rather than in month seven.
Common questions
Screen on artifacts rather than credentials. Ask to see the prompt set, the scoring methodology, and a sample baseline. Ask what percentage of scope is off-site and what they cannot promise. A provider running these programs answers all of it from existing documents.
Start with the prompt set, since it sorts most of the market immediately. Then engines tested, scoring methodology, off-site percentage, baseline timing, month one activities, contract metrics, stated limitations, and what happens if the fixes turn out to be finite.
Below roughly $1,500 per month, where the market openly describes offerings as repackaged SEO with an AI label. The hours that price buys cannot cover multi-engine measurement plus structural work plus any off-site citation supply. Prices far above the range should be itemized.
Yes, and it should be published. No industry-standard formula exists, and no independent cross-platform accuracy benchmark had been published as of August 2026. A provider claiming to use an industry standard is describing something that does not exist.
Citation frequency, citation accuracy, AI referral traffic in analytics, and share-of-voice movement against named competitors, each with named re-test dates. Rankings and generic traffic targets indicate an SEO scope with new labeling rather than a GEO program.
Not reliably. The category is young enough that size mostly signals overhead rather than capability. What distinguishes providers is whether they own a documented methodology and can produce a prompt set and baseline on request, which is independent of headcount.
Start here
The GEO program from Conspicuouz Creative Group publishes its methodology, its six scoring dimensions, what happens in each phase, and what gets measured at Day 45 and Day 90. Hold it against the nine questions above, then hold every other proposal against the same nine. On question nine, the answer is already published: when an audit finds the fixes are finite, we say so and recommend a project instead of a retainer.