Process · Three Phases
Ninety days is enough to prove a program is working. It is not enough to show what it produced. Six months is not padding, it is how long the slowest input takes.
By Izzy Gregorio · Updated August 2026 · 13 min read
In short
A 180-day GEO pilot runs in three phases: foundation days 1 to 45, authority days 46 to 120, amplification and measurement days 121 to 180, with re-tests at Day 45, Day 90, and Day 180 against a frozen prompt set recorded on Day 0. The length comes from the constraint rather than the contract.
Why 180 days
Three timelines govern this work and they are not equal.
Crawler access and indexation resolve in days. Structural editing takes effect at the next crawl. Third-party citation supply takes months, because it depends on publishers, reviewers, and roundup authors acting on their own calendars.
Market guidance puts first meaningful movement at 14 to 45 days and measurable gains at three to six months, with payback typically landing between six and twelve months. Those are vendor estimates rather than independent research, and vendor guidance describes contracts under 30 days as unable to produce measurable results.
Our operating position is that most programs take six to eight months to produce results worth reporting to a board. That is more conservative than the three to six month figure the market publishes, and it is stated deliberately.
A program that shows early citation movement at Day 45 and calls it a result has measured a crawler fix.
So the phases map to the constraint. Ninety days starts the slowest engine. One hundred and eighty days lets it run long enough to read.
Phase 1 | Days 1 to 45
Infrastructure before content. Everything measured later is measured against what gets recorded here.
Baseline prompt test. A frozen set of roughly 25 prompts across four engines, run logged out, raw results recorded and dated. This is Day 0.
Crawler access verification. GPTBot, OAI-SearchBot, ChatGPT-User, Google-Extended, PerplexityBot, ClaudeBot, Bingbot, CCBot. Plus CDN and firewall bot rules, which robots.txt will not reveal.
Indexation check across Google and Bing, since Bing's index remains one of several inputs feeding ChatGPT retrieval alongside OpenAI's own crawler and licensed publisher feeds.
Entity and listing consistency. Name, address, phone, hours, and descriptions made identical everywhere they appear.
Structured data audit and implementation. Organization sitewide, Article with a real linked author, FAQPage matching visible content.
Competitive citation gap against three to five named competitors.
Off-site citation supply planning begins in week five, earlier than most programs, because the placements take the longest to land.
What you provide: analytics and search console access, CMS access or a developer contact, your real competitor list, and the customer situations your buyers actually have.
Day 45 re-test. Same prompt set, same engines, logged out. Expect access issues resolved and early citation-rate movement on any pages already restructured. Do not expect share-of-model change.
A provider who has run these programs says that in advance rather than explaining it afterward.
Phase 2 | Days 46 to 120
Citation-worthy content on a foundation that can now be read. This is the longest phase and it carries most of the work.
Structural editing on the 10 to 15 highest-value existing pages, then extending across the site. Self-contained answers moved to the top, backward dependencies removed, headings rewritten as questions. This is editing rather than new production and it is the highest-return work in the pilot.
Cluster gap filling. The prompt types you are absent from get assets, in the format that answers them.
Proof density added. Research presented at KDD 2024 measured citation lift across roughly 10,000 queries: adding statistics produced the largest single lift at about 41 percent, citing external sources produced gains up to 115 percent for lower-ranked content, and adding quotations lifted visibility around 28 percent.
Off-site citation supply in motion. Listings corrected, directory and association coverage, comparison and roundup outreach, earned media pitched and placed.
First refresh cycle begins around day 110. Amsive found half of all AI-cited content is under 13 weeks old, and AirOps found pages unrefreshed for three months are more than three times as likely to lose citations entirely. The earliest phase-two work is already aging.
Day 90 re-test. The midpoint comparison. By now citation rate should be moving clearly and share of model should be showing early direction. This is the checkpoint where a program that is not working becomes visible, which is the entire reason it exists.
Day 90 is also the natural decision point. If the trend line is flat here with access and structure already fixed, the diagnosis was wrong and the second half should be re-scoped rather than continued.
Phase 3 | Days 121 to 180
Compounding, distribution, and the comparison that decides what happens next.
Earned placements from phase two land and accumulate. This is why the outreach started in week five.
Second refresh cycle runs. Every asset from phase two is now past its freshness window.
Original data published, if the program includes it. A number nobody else has is the only asset that reliably makes other people cite you by name, and six months is the first engagement length where publishing one and seeing it referenced fits inside the term.
Paid distribution only behind content that has proven itself organically. Putting budget behind an asset the engines have already declined is buying reach for a rejected source.
Day 180 re-test and the full before-and-after against Day 0.
What you should see by Day 180: share of model movement, a narrowed competitive citation gap against named competitors, accumulated third-party citations, and a documented comparison that answers whether the program worked without requiring anyone to take a position on it.
Translate it
| Phase | Your time commitment | What you approve | What breaks it |
|---|---|---|---|
| Foundation, days 1 to 45 | 4 to 6 hours total. Access provisioning, competitor list, situation mapping | The prompt set. Do not skip this, it defines everything measured afterward | Delayed access. The most common cause of a stalled phase one |
| Authority, days 46 to 120 | 2 to 4 hours per month. Content review and subject-matter input | Structural edits, new asset outlines, earned media angles | Slow content approval. Every week of review delay costs crawl cycles you cannot buy back |
| Amplify, days 121 to 180 | 1 to 2 hours per month. Placement and spend approvals | Original data publication, paid spend if any | Nothing much, if phases one and two ran on schedule |
The middle row is where most programs lose time and it sits entirely inside the client's control.
The contract
Five items. If they are not in the agreement, the pilot is not measurable.
The frozen prompt set, attached as an exhibit.
The Day 0 baseline, delivered before any production work begins.
Three named re-test dates: Day 45, Day 90, and Day 180.
A Day 90 decision point, written in, with the condition that triggers a re-scope.
The metrics: citation frequency, citation accuracy, AI referral traffic in analytics, and share-of-voice movement against named competitors.
A pilot without named re-test dates is a retainer with a trial label.
Item four is the one that protects you specifically. A six-month commitment without a mid-point off-ramp asks you to fund the second half before anyone knows whether the first half worked.
Set expectations honestly
Three things, stated in advance because discovering them in month five is expensive.
Category ownership. Semrush's 2026 study of 1,094 US categories found owners hold position 90.4 percent of the time month over month, which cuts both ways. Sticky positions are hard to take. Six months establishes presence and starts a claim. It does not unseat an entrenched owner.
A clean attribution line to revenue. AI referral traffic undercounts badly, since a buyer who reads about you in an answer and later types your name arrives as direct or branded search. Expect strong directional evidence at Day 180, not a closed loop.
Payback, necessarily. Market guidance puts payback at six to twelve months. Day 180 is the front edge of that window, not the middle of it. A program that is working may still be pre-payback when the term ends, and knowing that going in is the difference between a renewal conversation and an argument.
One question worth sitting with. Six months from now someone will ask whether this worked. The only thing that determines whether that question has an answer is whether a number was recorded before the work started.
So before you evaluate any proposal on price, ask three things: when the baseline gets recorded, what happens at the midpoint if it is flat, and who keeps the prompt set when the term ends.
Common questions
Three phases. Foundation days 1 to 45 covering baseline testing, crawler access, entity consistency, and structured data. Authority days 46 to 120 covering structural editing, cluster gaps, and off-site citation work. Amplify days 121 to 180 covering placements, refresh cycles, and the Day 180 comparison.
Because the largest lever runs on other people's schedules. Roughly 82 percent of AI citations come from earned media, and publishers, reviewers, and roundup authors publish on their own calendars. Crawler fixes resolve in days and structural edits at the next crawl, but citation supply cannot be compressed.
A documented Day 0 baseline on a frozen prompt set, verified crawler access, confirmed indexation, corrected listings, and a prioritized fix list. Not citations. Citation movement is not a month one output, and a provider promising it is describing technical fixes in dramatic language.
Market guidance puts first meaningful movement at 14 to 45 days, usually as citation-rate improvement on restructured pages rather than share-of-model growth. Those are vendor estimates rather than research. Day 45 is the first fair comparison. Day 90 is where a program that is not working becomes visible.
Analytics and search console access, CMS access or a developer contact, a real competitor list, your buyers' actual situations, and timely content approvals. Roughly four to six hours in phase one, then two to four hours monthly. Delayed access is the most common cause of a stalled start.
The frozen prompt set as an exhibit, the Day 0 baseline delivered before production work, three named re-test dates at Day 45, 90, and 180, a written Day 90 decision point with a re-scope condition, and specific metrics including citation frequency, accuracy, AI referral traffic, and share-of-voice movement against named competitors.
You should be able to, and the contract should say so. The Day 90 re-test exists as a decision point. If the trend is flat with access and structure already fixed, the diagnosis was wrong, and the second half should be re-scoped rather than continued on momentum.
Start here
Three phases, three named re-test dates, a written midpoint decision point, and a before-and-after comparison at Day 180 from the GEO program at Conspicuouz Creative Group. The prompt set is yours, so the measurement survives the engagement.