I tested three AI visibility tools on the same brand. They agreed on the result, not the evidence.

Steve New

I wanted a baseline for stevenew.com, so I put the same eight commercial questions through three AI visibility platforms: SE Ranking, Otterly and MentionScore. The questions covered subjects I actually write and consult on, including website rebuilds, CRO audits, A/B testing, price objections and AI visibility itself.

Across 104 sampled prompt-engine answers, Steve New appeared zero times. For this defined question set and measurement period, that makes the broad conclusion fairly convincing: my visibility was effectively zero.

The evidence underneath that conclusion was much less consistent: after normalising the source URLs, the three systems collectively returned 407 different cited pages, but only 19 appeared in all three. There were 273 different source domains, of which only 29 appeared across all three systems.

Even under the shared label “ChatGPT”, Otterly returned 26 unique cited pages across the eight questions and MentionScore returned 42. Only 10 appeared in both sets.

The comparison does not tell me that one tool was right and another was wrong. These were different measurement systems, not three instruments observing an identical set of underlying AI runs. What interested me was that the coarse conclusion survived those differences much better than the detailed evidence did.

An AI visibility score is not a direct reading of some fixed property called visibility. It is the output of a measurement system.

What I compared

The same eight questions formed the shared comparison set, using US English and as close a measurement window as practical.

ToolQuestionsAI surfaces in my runEvidence available
SE Ranking8Google AI Overview, Google AI Mode, ChatGPT, Gemini, PerplexityRankings, cached answers, source pages and domains
Otterly8 controlled questionsChatGPT, Google, Perplexity, CopilotFull answers, citations, prompts, source pages and domains
MentionScore8ChatGPT, Gemini, Google AI Mode, Google AI OverviewEntity/visibility reporting, mentions, citations and source data

This was not a laboratory-clean comparison.

Otterly required 15 prompts during onboarding, so I added seven filler questions and excluded them from this analysis. MentionScore treated Steve New primarily as a person/entity rather than as a pure domain project. The engine sets also differed, and a label such as “ChatGPT” does not establish that two vendors are using the same model, search configuration, retrieval layer or endpoint.

For the source comparison, I stripped tracking parameters and fragments while preserving the underlying page identity, then deduplicated the resulting URLs.

So the experiment tells me what three different measurement systems observed. It does not isolate how much of the disagreement was caused by vendor methodology itself.

What stayed stable, and what didn't

The main results were:

FindingResult
Sampled prompt-engine answers104
Steve New appearances0
Distinct source pages407
Pages appearing in all three tools19
Distinct source domains273
Domains appearing in all three29
Otterly “ChatGPT” pages26
MentionScore “ChatGPT” pages42
Shared pages under the “ChatGPT” label10

I am comfortable with the coarse conclusion that Steve New was effectively absent from this defined sample. Eight questions are not the whole market, so I would not generalise that into a claim that the brand simply “doesn't exist in AI search”. But zero appearances across 104 samples from several measurement systems is enough for me to treat weak visibility as a real baseline rather than demand another decimal place.

I would put much less confidence in an exact competitor percentage, a source ranking or a small movement in share of voice.

The source landscape makes the difference clear. Across the three systems, only 19 of the 407 distinct cited pages appeared in all three. The domain picture was somewhat more stable, which is unsurprising because the same site can recur while the precise article changes.

The apparently comparable “ChatGPT” results deserve particular caution. Otterly and MentionScore were given the same eight questions, but their source sets overlapped on only 10 pages. That does not establish that they sent identical requests to an identical ChatGPT configuration and received conflicting answers. It establishes something more useful: seeing the same surface label on two dashboards does not guarantee that the measurements underneath it are equivalent.

If one report suddenly tells me that a particular page, domain or competitor has become strategically important, I want more evidence before reorganising a content programme around it.

What sits between the question set and the score

Before a dashboard displays a visibility score, several choices and processes have already shaped it.

Someone has first decided what counts as the market being measured: the prompts, language, geography, AI surfaces and competitors. Ten broad informational questions are measuring something different from fifty late-stage buyer questions.

The AI systems then produce particular answers. They decide whether and how to search, which sources to retrieve, what to mention and what to cite, and those outputs are not perfectly deterministic.

Finally, the visibility platform has to interpret what came back. It resolves the brand as an entity, identifies mentions and citations, constructs competitors and aggregates multiple observations into a score or share-of-voice measure.

I saw a literal example of the last step during the Otterly setup. Its initial competitor detection interpreted “Steve New” through a shoe and fashion context and suggested Steve Madden alongside several footwear retailers.

For an obscure personal brand with an ambiguous name, the inference was understandable. But had I accepted it, the resulting competitor analysis would still have appeared in a polished dashboard.

The precision of an interface is not the same thing as precision in the underlying measurement.

Some variation can also arise before the visibility vendor interprets anything.

A September 2026 study published by AI-search vendor Prefer asked the same 80 questions three times each across ChatGPT, Gemini, Claude and Perplexity. Among pairs of runs where both answers cited sources, two runs of the same question shared on average 37% of cited sites on ChatGPT and 35% on Gemini. Claude was more stable at 68%, while Perplexity reached 96%. The runs were made seconds apart through API access with web search requested.

That is vendor-produced research rather than independent academic evidence, and API behaviour should not automatically be treated as identical to the consumer applications. The method and dataset are published, though, and the result is directly relevant to the measurement problem.

A 2026 arXiv preprint by Ronald Sielinski reaches a related conclusion from repeated sampling across generative-search platforms. It treats citation visibility as a sample estimate from a distribution of possible responses rather than as a fixed quantity, and reports substantial variation in citation shares and rankings across repeated observations.

This means I cannot look at my own 407-page source universe and assign the differences cleanly to the vendors. Some may come from study design or vendor implementation, some from the underlying models and retrieval systems, some from timing, and some from ordinary run-to-run variation.

The defensible conclusion is narrower: what appears in an AI visibility dashboard is conditional on both the measurement setup and the particular AI observations the tool collected.

The finer the claim, the more evidence I want

If my brand appears in zero of 104 commercially relevant samples across several different measurement systems, I am comfortable saying visibility is weak for that defined set.

If one source appears in three answers rather than two, or a share-of-voice score moves from 17.3% to 19.1%, I want to know more. Were the same prompts used? The same engines? How many observations sit behind the number? Was the entity resolved in the same way? Did the underlying answers change, or only the way they were aggregated?

A decimal point does not manufacture certainty.

Polished dashboards can make a broad absence, a detailed source ranking and a small score movement look similarly precise even though they support very different strengths of inference.

That matters because the action attached to the finding may also be very different. I would accept a persistent broad absence much sooner than I would reorganise a content strategy because one source moved from seventh to fourth in a report.

How I would use an AI visibility tool

I would still use one, but I would treat the measurement setup as part of interpreting the result.

I would start by defining what I am trying to observe and choosing questions that represent real commercial territory rather than allowing a generated prompt set to become the market by default. If I want to monitor change, I would keep the setup reasonably stable. Changing the prompts, engines, entity definition and vendor at the same time changes the measuring instrument as well as whatever I hoped to measure.

I would also keep the raw evidence behind important movements. If a score changes sharply, I want to inspect the answers, citations and prompts before inventing an explanation for it. The finer or more consequential the finding becomes, the more willing I am to repeat or triangulate it.

I would not compare a score of 42 in one product with 57 in another as though the second platform had discovered fifteen additional units of the same underlying thing. The systems are not necessarily estimating an identical quantity in an identical way.

Even a robust visibility observation does not explain why that state exists. That is a separate diagnostic problem.

For this experiment, the measurement conclusion is enough. Three systems agreed that Steve New was effectively absent from this defined question set, while much of the detailed evidence underneath that result varied across the measurement systems and sampled observations.

I would therefore use AI visibility tools as deliberately defined measurement systems, not as interchangeable maps of a fixed market.

Sources