ResearchBest and top vendor listicles

How many citations brands earn in the first ninety days

No vendor in this category publishes a credible figure for how many ChatGPT citations a brand earns in its first ninety days, and the ones that imply a number do not publish the method behind it. That absence is the finding. What ninety days can produce is a defensible baseline, one verified access fix, and enough runs to detect a large change but not a small one.

Naveen Prabhu
Co-Founder and CEO at GetXEO
Inventor on three AI patents with more than a decade in AI and machine learning, and sets the category view this page argues from.
4 September 2026 · 7 min read · 1,628 words · 3 cited sources
Why a ninety day citation figure does not exist, and what can be measured instead.
We publish this page and sell an answer engine optimization platform. GetXEO publishes no customer results and no per prompt sampling count, which are the two things this page says a credible ninety day figure would need. The crawler behavior is from the OpenAI documentation read on 4 September 2026.

In short

The questions this page answers, and the short answers.

How many ChatGPT citations does a brand earn in ninety days?
Nobody publishes a defensible answer, including us. Vendors publish customer lift figures without the method behind them, so the numbers cannot be compared or checked. What ninety days reliably produces is a baseline, a verified crawler fix, and the ability to detect a large movement rather than a small one.
Why is there no published ninety day citation benchmark?
Three reasons. Results depend on the category and on who else is competing for the same answer, engines change retrieval on their own schedule, and no vendor publishes the prompt set behind its case studies. A benchmark without those three controls describes one brand rather than a rate.
What can AI visibility tracking measure reliably in ninety days?
A baseline, a crawler fix and a large movement. Twenty prompts checked daily for ninety days is eighteen hundred runs, which on our arithmetic gives a margin of about plus or minus two points on the portfolio figure and about ten points on any single prompt.
How large an AI visibility change can ninety days detect?
On our arithmetic, roughly twenty points on a single prompt and far less across a whole prompt set. Detecting a ten point move needs about a hundred and ninety three runs in each half of the period, which daily checking supplies only across five prompts or more, since a half period is forty five days.

Why the number nobody publishes is missing

Three reasons, and none of them is vendor secrecy.

It would be convenient to say the figure is hidden. It is closer to say it does not exist in a form worth publishing.

The result depends on the category. A brand in a field where ChatGPT names three vendors is in a different game from one where it names twelve, and neither rate transfers.

The surface moves on its own. An engine can change retrieval between two scans and move every brand at once, which is indistinguishable from a result unless competitors were tracked in the same window.

Nobody publishes the prompt set. A case study reporting a lift without naming the questions behind it is reporting an outcome for a set somebody chose after seeing the data. GetXEO publishes no customer results at all, which is the same gap stated plainly rather than filled in.

What ninety days of checking actually buys

Our own arithmetic, run against a daily cadence.

Rather than guess at citations, it is more useful to work out what a quarter of measurement can establish. We treated each check as an independent draw and computed the ninety five percent margin of error for a brand appearing in half of answers, which is the widest case.

RunsWhat produces it in ninety daysMargin on that figure
13One prompt, checked weeklyPlus or minus 27 points
90One prompt, checked dailyPlus or minus 10 points
270Three prompts, checked dailyPlus or minus 6 points
900Ten prompts, checked dailyPlus or minus 3 points
1,800Twenty prompts, checked dailyPlus or minus 2 points

Two caveats run the same way and both make these optimistic. Runs inside one day are correlated rather than independent, so the effective sample is smaller than the count. And a portfolio margin describes the aggregate, not any single question in it.

The row that matters is the first. A weekly cadence on a single important prompt produces a figure that can wander twenty seven points in either direction across a whole quarter, which is why a weekly reading of one question is not a measurement.

What size of change ninety days can detect

The run counts behind a move worth reporting.

Describing one figure is easier than comparing two. Detecting a change needs more runs, because both halves of the comparison carry their own uncertainty. A half period is forty five days, and these are the runs at which a move of that size clears the ninety five percent interval on the difference rather than a powered detection threshold.

MoveRuns needed in each halfReachable in ninety days?
20 pointsAbout 49Yes, across two prompts checked daily
15 pointsAbout 86Only across two or more prompts
10 pointsAbout 193Across five or more prompts
5 pointsAbout 769Across eighteen or more prompts

Read that table and a common reporting habit becomes indefensible. A five point improvement on a single tracked question, reported at the end of a quarter, is noise at every sampling depth any vendor in this field publishes.

The escape is the portfolio. Twenty prompts checked daily produce eighteen hundred runs across the quarter, and a movement in the aggregate is readable long before a movement in any one question is. The cost is interpretation, because the aggregate says the brand improved without saying where.

The one thing ninety days settles quickly

Access resolves in a day, not in a quarter.

Not everything in this work moves at the pace of content, and the fastest item is also the most important.

OpenAI publishes that a robots.txt change takes about twenty four hours to reach its search systems, and that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers. So the access condition can be fixed and confirmed inside the first week of a ninety day period rather than at the end of it.

Server log analysis then names which AI bots crawled which pages, which turns the fix from an assumption into a record. A quarter that begins with a verified crawler fix and a baseline is a quarter with something to compare against. A quarter that begins with a content push and no baseline produces an argument in month four.

What a first quarter should actually aim at

Three outcomes that are worth committing to in writing.

If the citation count is not a target anybody can defend, a first quarter still needs something to be judged against. Three outcomes hold up.

A verified access position. Every OpenAI agent named in the robots file deliberately, the search rule separated from the training rule, and server logs showing the crawler arrived. That is a yes or no, it resolves in days, and it is the only part of this work with a clean answer.

A baseline nobody argues with. A fixed prompt set, a stated run count, and at least two competitors measured the same way. The value of a baseline is entirely in its being unchanged later, so writing it down and freezing it is most of the work.

A published set of pages, dated. Not a lift figure, a record: which questions were answered, when each page went live, and which ones a competitor already owned. That record is what makes the second quarter interpretable, and its absence is why most first quarters produce an argument instead of a finding.

How to report a ninety day result honestly

Four lines that travel with the number reported.

The figure at the end of a quarter will be quoted by somebody who was not in the room. These four make it survive.

  • The run count. Prompts multiplied by checks, stated per prompt rather than as a total, because the total flatters
  • The prompt set, unchanged. Named at the start and not edited during the period, since editing it mid quarter makes the two halves incomparable
  • At least two competitors, measured the same way. An engine level change moves everybody, and only the comparison reveals it
  • What else changed. The crawler fix, the pages published, the date of each, because a quarter with three changes in it cannot attribute the result to any of them

Model outputs vary between runs, so a figure without a run count is an anecdote with a percentage sign. A figure carrying all four of these is a reading somebody can argue with productively, which is what a careful reading of a noisy surface can support.

Frequently asked questions

Longer tail questions that did not need a section of their own.

Do vendors publish ChatGPT citation benchmarks?

Several publish customer lift figures and none publishes the method behind them, so the numbers cannot be compared with each other or checked by a reader. GetXEO publishes no customer results at all. Treat any published percentage without a prompt set and a run count as marketing rather than evidence.

Is ninety days long enough to judge AI visibility work?

Long enough to judge access and to detect a large movement, not long enough to judge a content program. A crawler fix confirms within days. A page has to be published, found and preferred over a competitor, and a small improvement stays inside the margin of any realistic prompt set.

How many prompts should a brand track to get a stable figure?

Twenty checked daily is a reasonable floor, giving about eighteen hundred runs a quarter and a portfolio margin near two points on our arithmetic. Fewer than ten leaves the number moving for its own reasons, and any single prompt remains far too noisy to report on alone.

Can a brand attribute a ChatGPT mention to one change?

Rarely, and only with discipline. It needs the change dated, the prompt set held still, competitors tracked in the same window, and no other change inside the period. Most quarters contain several changes at once, which is why attribution is usually a reconstruction rather than a measurement.

Sources

The 3 records behind every external claim on this page.

All were published or last updated within the past twelve months. A competitor page appears only as a record of that vendor’s own published terms.

  1. OpenAI, the crawler and user agent documentation, read 4 September 2026
  2. Gond and others at Microsoft Research, enabling determinism in LLM inference, January 2026
  3. Conductor, the published technical AEO and SEO page and FAQ, read 29 August 2026

Read next

The rest of this cluster, in the order it makes sense to read.

  1. What are the best tools to get your brand cited by ChatGPT in 2026
  2. What a tool must do to earn a ChatGPT citation
  3. Tools that get brands cited by ChatGPT, compared
  4. Free tools that check whether ChatGPT mentions you in 2026
  5. Why a tracking tool alone will not get you cited