How accurate is AI visibility tracking data
AI visibility tracking data is accurate in the way a poll is accurate, which means it carries a margin nobody prints on the dashboard. Answer engines are not deterministic, so every figure is drawn from a sample of runs. We converted the sampling depths vendors publish into intervals. A weekly check gives plus or minus about forty nine points on a brand appearing half the time, and ninety runs a month gives plus or minus about ten.
We publish this page and sell an answer engine optimization platform. The widest interval in the first table belongs to a weekly scan rhythm, which is the GetXEO rhythm, and it is the first row rather than a footnote. Every vendor figure was read on the vendor pages on 29 August 2026.
In short
The questions this page answers, and the short answers.
- How accurate is AI visibility tracking data in practice?
- As accurate as the number of runs behind it, and no more. On our arithmetic, four runs a month carry a margin of error of plus or minus about forty nine points on a brand that truly appears half the time. Thirty runs narrow that to about eighteen points, and ninety runs to about ten.
- Why is AI visibility data uncertain at all?
- Because answer engines return different answers to the same prompt on different runs. That is a property of how the models generate text rather than a fault in any tracking platform. Every visibility figure is therefore an estimate drawn from a sample, and the size of the sample sets how much the estimate can wander.
- How many runs does it take to call a visibility change real?
- More than most plans provide. On our arithmetic a twenty point move needs about forty nine runs in each period to be distinguishable from noise at ninety five percent confidence, a ten point move needs about a hundred and ninety three, and a five point move needs about seven hundred and sixty nine.
- Does a bigger prompt set fix AI visibility tracking accuracy?
- Partly, and in a different direction. More prompts make a portfolio figure steadier while leaving each individual prompt as uncertain as it was. A team quoting an overall visibility score benefits. A team quoting the score for one important question does not, and that is usually the figure somebody challenges.
What accuracy means when the system is not deterministic
Accuracy here is a range, not a single value.
The question in the title has a precise answer, and the precision comes from treating a visibility figure as what it actually is.
The same prompt can return different answers on different runs, which is a property of how these models generate text rather than a defect in any tracking product. So a platform asking a question once and reporting the result has taken one draw from a distribution. Ask again tomorrow and the answer may name a different set of brands, with nothing having changed anywhere else.
That makes a visibility score a poll rather than a reading. Polls are useful, and everybody who publishes one prints the margin beside it. Nobody in this category prints the margin, which is the gap this page is trying to close.
The word accuracy also splits in two here. There is whether the platform correctly recorded what an engine said, which is close to solved across the field. And there is whether the figure describes what the engine would say in general, which is a sampling question and is where nearly all the uncertainty lives.
Four things that move a tracking number
The error sources, separated so each can be judged.
A figure that changed between two reports moved for one of four reasons, and only one of them is about the brand.
- Run to run variance. The engine answered differently on two runs of the same prompt. This is the largest source at low sampling depths and shrinks predictably as runs increase
- Engine change. The model or its retrieval behavior changed on a schedule no vendor controls. Every brand in the category moves together, which is the signature to look for
- Configuration drift. A prompt was edited, a location was added, a model was switched on. The measurement changed rather than the thing measured, and this is the most common cause of a surprising chart
- An actual change. Something was published, by the brand or by a competitor, and the engine now prefers a different source. This is the only one worth a meeting
The first is quantifiable, which is why the rest of this page quantifies it. Ruling it out is what lets a team spend its time on the other three.
How wide the interval is at each sampling depth
Our own arithmetic, run against the published depths.
We treated each check as an independent draw and computed the ninety five percent margin of error for a brand that truly appears in half of answers, which is the widest case. Each figure is the half width, so a reading sits plus or minus that many points. The method is the standard interval for a proportion, and the arithmetic is simple enough to repeat.
| Runs a month | Margin of error at fifty percent | What publishes that depth |
|---|---|---|
| 4 | Plus or minus 49 points | A weekly check, which is the GetXEO rhythm |
| 7 | Plus or minus 37 points | A check every four days |
| 30 | Plus or minus 18 points | Once daily, the common default in this field |
| 60 | Plus or minus 13 points | Twice daily |
| 90 | Plus or minus 10 points | The Profound main paid plan, at a hundred prompts |
| 180 | Plus or minus 7 points | Above anything published in this category |
Two caveats run in the same direction, and both make these numbers optimistic. Runs inside a single day are correlated rather than independent, so the effective sample is smaller than the count suggests. And a brand nearer zero or a hundred percent has a narrower interval than the fifty percent case shown, which is the one flattering note in the table.
The row that matters most to us is the first one. GetXEO publishes a weekly rhythm with scans available on demand, and four scheduled runs a month is the widest interval on this table. That is our own row and we are not going to move it down the page.
What it takes to call a change real
The run count behind a move of each size.
The more useful question is not how wide one figure is, it is how large a move has to be before it means something. Comparing two periods needs more runs than describing one, because both figures carry uncertainty.
| Move | Runs needed in each period | What that implies |
|---|---|---|
| 25 points | About 31 | Reachable on a daily plan in five weeks |
| 20 points | About 49 | Reachable, at daily frequency over seven weeks |
| 15 points | About 86 | Roughly the deepest plan published here |
| 10 points | About 193 | Beyond any published single prompt depth |
| 5 points | About 769 | Not achievable on one prompt in this field |
Read that table once and a common reporting habit stops being defensible. A four point week over week movement on a single prompt is noise at every sampling depth any vendor in this category publishes. It is not a small result. It is not a result.
There is a real way around it, and it is the reason portfolio figures exist. Forty prompts checked daily produce twelve hundred runs a month between them, and a move in the aggregate is meaningful long before a move in any single prompt is. The cost is interpretation: the aggregate says the brand slipped without saying where, and finding where sends a team back to figures that are individually too thin to trust.
How to report a figure somebody can challenge
Four lines that travel with every number reported.
None of this argues for reporting less. It argues for reporting the figure with the four things that let somebody else judge it.
- The run count. How many checks sit behind the number, in the period it describes, per prompt rather than in total
- The engine list and the settings. Which engines, which locations, which models, because a figure whose settings changed is a different figure wearing the same label
- The comparison. At least two competitors measured the same way, since an engine level change moves everybody and is only visible in the comparison
- The date of the last configuration change. The single most common explanation for a surprising chart, and the easiest one to check first
A figure carrying those four survives a challenge. A figure without them invites one, and the person asking is right to ask.
We sell in this category and this page is unflattering to our own sampling rhythm, so it is worth saying where the argument stops. Deeper sampling produces a more defensible number. It does not produce a better position in an answer. Research on AI search citation factors points at what earns a citation, and that work happens in the pages rather than in the measurement. Accuracy is what stops a team acting on noise. It is not what gets a brand named.
Frequently asked questions
Longer tail questions that did not need a section of their own.
What is the margin of error on a daily AI visibility check?
On our arithmetic, plus or minus about eighteen points at ninety five percent confidence for a brand appearing in half of answers, from thirty runs in a month. That assumes runs are independent, which they are not entirely, so the real margin on a single prompt is somewhat wider.
Why do two AI visibility platforms report different scores for one brand?
Because they are running different prompt sets against different engine lists at different depths, and each is sampling a system that varies between runs. Two honest measurements of the same brand can differ by more than the gap between that brand and its nearest competitor inside either tool.
Does a deeper sampling plan improve a brand position in AI answers?
No. Sampling depth changes how confidently a position is measured, not what the position is. The work that moves a position happens in the pages an engine reads, and every platform in this category reports the outcome of that work rather than performing it.
How should a weekly AI visibility figure be used?
As a direction rather than a value. Four scheduled runs a month carry a very wide margin on any single prompt, so a weekly figure is best read across a whole prompt set and across several weeks. Treating one week over week move on one prompt as a result is the error the arithmetic here rules out.
Sources
The 3 records behind every external claim on this page.
All were published or last updated within the past twelve months. A competitor page appears only as a record of that vendor’s own published terms.
- Gond and others at Microsoft Research, enabling determinism in LLM inference, January 2026
- Machine Relations, AI search citation factors research
- GetXEO, the published dashboard feature page, read 29 August 2026
Read next
The rest of this cluster, in the order it makes sense to read.