Four AI Engines Agreed on the Same Citation Only 1.7% of the Time

Four AI Engines Agreed on the Same Citation Only 1.7% of the Time
Sponsored

AI search visibility is often discussed as if ChatGPT, Claude, Gemini and Perplexity were four interfaces into roughly the same answer ecosystem. A new practitioner study suggests the opposite: when the models decide which sources deserve citations, agreement can be exceptionally rare.

Bill Widmer tracked 13,184 citations across 1,765 AI-generated answers for an Orbit Media Studios study published on September 3. The experiment covered 72 prompts related to three B2B brands and ran those questions repeatedly through ChatGPT, Claude, Gemini and Perplexity.

Across 1,792 query-and-domain combinations, all four engines cited the same domain for the same question only 30 times.

That is 1.7%.

The finding does not establish a universal probability that AI engines will disagree on any arbitrary query. The sample is narrow, commercially oriented and deliberately described by its author as directional rather than statistical proof. But it provides a useful warning for brands trying to turn AI visibility into a single ranking number: different models can construct answers from substantially different parts of the web.

The study tracked 72 buyer-oriented prompts across three B2B brands

Widmer built his own LLM visibility tracker with Claude Code and used it to monitor three sites: his own billwidmer.com, Amazon agency Evolve Ad Agency and SEO platform Semrush.

Each brand contributed between 12 and 40 prompts, for 72 questions overall.

The prompts were designed around real commercial questions rather than conventional short keywords. Examples included queries about the best all-in-one SEO platform for marketing teams and whether paying an agency to manage Amazon is worthwhile.

A script ran the questions weekly through the four AI systems, recorded the generated answers, extracted cited URLs and tracked whether the relevant brand was mentioned.

The dataset analyzed in the article runs through August 23, 2026.

This is important context. The study is not a random sample of everything people ask AI. It is a focused B2B experiment involving three brands and a fixed panel of commercially relevant questions.

Only 30 combinations produced a citation shared by all four engines

The headline result comes from 1,792 query-and-domain combinations.

For each question, the analysis looked at the domains the models chose to cite. A domain counted as four-engine agreement only when ChatGPT, Claude, Gemini and Perplexity all cited it for the same question.

That happened 30 times, or 1.7% of the measured combinations.

Even pairwise agreement could be low. Widmer reports that ChatGPT and Perplexity shared 7.6% of their cited domains on one brand’s question set and 5.4% on another.

The result challenges a convenient assumption behind generative engine optimization: that there is a stable collection of universally trusted sources that brands simply need to enter.

At least in this dataset, source selection was far more fragmented.

Each model displayed a different citation appetite

The engines did not merely select different domains. They also cited dramatically different numbers of sources.

Perplexity was the most citation-heavy system in the study, averaging 19.2 sources per answer. Gemini averaged 8.1, ChatGPT 4.5 and Claude 3.6.

That alone makes raw citation counts dangerous as a cross-platform KPI.

If one engine routinely attaches roughly five times as many sources as another, a brand can accumulate more citation opportunities there without necessarily being more influential in the answer.

The source preferences also differed. Widmer found Perplexity frequently citing LinkedIn, YouTube and Reddit across the full dataset. Gemini leaned toward Reddit and list-style roundups. ChatGPT skewed more toward primary sources such as Google documentation, academic papers and vendor help or pricing pages, while Claude frequently used review platforms and curated lists.

These are observations from this particular B2B prompt panel, not permanent descriptions of how the products retrieve information.

Reddit was almost absent in the latest B2B run

One result is especially relevant to the widespread advice that brands seeking AI visibility should simply “be on Reddit.”

In the study’s latest run of 72 B2B buyer questions, Reddit accounted for 0% of Claude citations, 0% of ChatGPT citations, 3% of Perplexity citations and 1% of Gemini citations.

Widmer says Claude did not cite Reddit at all across hundreds of answers in his broader dataset.

That does not mean Reddit is unimportant to AI systems generally.

The prompts were B2B questions tied to three particular brands. Consumer products, travel, entertainment, health, hobbies or other categories could produce a completely different source mix. AI providers can also change retrieval behavior rapidly.

The more defensible lesson is that source strategy should be based on the sources actually appearing for a brand’s category and questions rather than on universal GEO checklists.

Google’s Top 10 and AI citations barely matched at the exact-URL level

Widmer also compared the AI citations with conventional Google results.

Where a conversational prompt could be mapped to a similar or identical Google keyword search, he collected the organic Top 10 during the same week and measured overlap.

At the domain level, between 10% and 30% of domains cited by the AI models also appeared in Google’s Top 10 for the corresponding search, depending on the model and question set.

The exact-page overlap was substantially lower.

For Gemini and ChatGPT, only roughly 1% to 5% of cited URLs were also exact URLs in Google’s Top 10. Perplexity aligned more closely with Google, but even there the exact-URL overlap was only 13% to 23%.

This is not evidence that Google rankings are irrelevant to AI visibility.

It shows that an AI system can retrieve or cite a different page from the one Google ranks on page one for a roughly equivalent information need.

Domain overlap and URL overlap answer different questions

The distinction between domain and URL is crucial.

Suppose an AI answer cites a Semrush research article while Google’s Top 10 contains a different Semrush page.

At the domain level, the systems agree: both surface semrush.com.

At the exact-URL level, they disagree.

That difference can explain why domain-level overlap reaches 10% to 30% while exact-page overlap falls into the low single digits for some models.

For brands, domain visibility may therefore be more stable than individual-page visibility. But for publishers deciding which article to optimize, URL-level selection still matters.

A dashboard that reports only domain mentions can conceal substantial page-level churn.

The models also disagreed with themselves

Cross-engine disagreement was only part of the volatility.

Widmer ran the same 72 questions through the same models twice, two days apart, then compared the citations.

Claude retained only 30% of the URLs it had cited previously. Gemini retained 38% of its domains. Perplexity was more stable at 65%.

The study does not provide a comparable ChatGPT retention figure in that section, so the three reported percentages should not be extrapolated to the fourth engine.

Still, the test illustrates a central problem in AI visibility measurement: one answer is not necessarily reproducible even when the question and model remain nominally the same.

A source can disappear without the publisher changing anything.

A single prompt run is closer to a sample than a ranking

Traditional rank tracking became useful because search positions, while variable, can be observed repeatedly against a relatively clear SERP.

Generative answers behave differently.

The response can change, the source set can change and the number of citations can change between runs. Retrieval can also depend on model configuration, product surface and provider updates.

That makes a one-time screenshot of “we rank in ChatGPT” weak evidence of durable visibility.

Widmer’s study argues for repeated measurement. His own framing is useful: a single answer is weather, while a tracked rate across weeks is closer to climate.

For GEO dashboards, that means persistence may matter as much as acquisition.

Mentions and citations are not the same metric

The dataset also separates brand mentions from citations.

A mention occurs when an AI answer names or recommends a brand. A citation occurs when it links to a page from that brand as a source.

Across 1,052 query runs examined for this relationship, only 17 cited a brand page without also mentioning the brand. The reverse was much more common: models often mentioned brands without linking to them.

The gap varied sharply by engine.

Claude mentioned the relevant brands in 70% of answers but cited them in 40%. Gemini was 72% versus 41%. Perplexity was 67% versus 44%.

ChatGPT had the narrowest gap: 56% mention rate and 47% citation rate.

That means citation share alone can understate brand visibility, particularly on systems willing to recommend a company without sourcing the recommendation from the company’s own site.

Famous brands may create an additional measurement problem

The study observed a particularly large mention-to-citation gap for Semrush.

Widmer proposes a possible explanation: established brands may already be strongly represented in model training data, allowing an AI system to mention them while retrieving third-party sources for supporting information. Less-known brands, by contrast, might be mentioned primarily when retrieval actually finds their sites.

This is an interesting hypothesis, not a demonstrated mechanism.

Widmer acknowledges confounders, including differences in the amount of content each brand has and whether a brand name was present in a prompt.

The observation is still useful because it highlights a measurement problem. A zero-citation result may mean something different for a famous company that is repeatedly recommended without links than for an unknown company that is neither mentioned nor cited.

AI visibility measurement therefore needs at least two dimensions: whether the brand enters the answer and whether its own properties are used as evidence.

The Claude setup limits direct platform comparisons

The study is unusually transparent about an implementation caveat that should not be buried.

Widmer says the Claude data came from an agent-with-search setup rather than the normal Claude interface or raw API configuration. He explicitly notes that its citation behavior may therefore differ from claude.ai.

The other three engines were run via APIs, with ChatGPT configured to force web search.

Gemini introduced another data limitation: it sometimes exposed only the cited domain rather than the exact page.

Those differences mean the experiment is not a perfectly controlled four-product benchmark.

It is better understood as a practical monitoring setup showing what happened under four specific retrieval configurations.

Conversational prompts do not map perfectly to Google keywords

The Google comparison has another unavoidable limitation.

A conversational buyer question is not always equivalent to a short search keyword.

Mapping “is it worth paying an agency to manage Amazon?” to a Google query requires an analytical choice about which traditional search best represents the same intent.

Even when wording is similar, Google and an AI assistant may be solving different versions of the information need.

That does not invalidate the overlap analysis. It does mean the 1% to 23% exact-URL figures should be treated as directional comparisons rather than proof that AI systems systematically ignore Google’s rankings.

The 1.7% result argues against one universal AI citation playbook

If all four engines routinely cited the same sources, GEO strategy would be relatively simple.

Brands could identify the common gatekeepers, win inclusion on those sites and expect visibility to propagate across models.

The Orbit Media data suggests a more fragmented environment.

Perplexity may retrieve many sources while Claude retrieves few. One model may favor primary documentation while another leans toward review platforms. A page can appear in Google’s Top 10 and never be selected by an AI system, while another page outside that Top 10 can become a citation.

That makes “rank in AI” an increasingly misleading phrase.

There may be no single AI ranking to win.

But different engines do not necessarily require four separate marketing strategies

Interestingly, Widmer does not conclude that brands need completely different strategies for every model.

His recommendation is closer to the opposite.

Rather than chasing a tactical trick for ChatGPT, another for Gemini and another for Perplexity, he argues for strong underlying marketing assets: genuine expertise, consistent information about the entity and credible third-party validation.

The distribution can differ even when the underlying work is shared.

A useful practical approach is to track real buyer questions repeatedly, identify the third-party sources that recur in a specific category and measure mentions and citations separately for each model.

That is less satisfying than a universal “AI ranking factors” checklist, but it fits the volatility the study actually observed.

13,184 citations still make this a directional B2B study, not a universal law

The headline sample size is substantial, but it should not be confused with broad population coverage.

The 13,184 citations came from 72 prompts around only three B2B brands. The questions were selected for those businesses, not randomly sampled from global AI usage.

The providers were accessed through different technical configurations, and citation extraction itself varies by platform. The Google comparison required mapping conversational questions to traditional queries. Repeatability testing covered a limited number of runs.

Widmer explicitly describes the project as directional rather than statistical proof.

That transparency strengthens the useful conclusion: this is evidence of substantial citation fragmentation within one monitored B2B environment, not a claim that every AI query in every category has exactly a 1.7% four-engine agreement rate.

AI visibility should be measured as a distribution, not a position

The most important lesson from the study may be methodological.

SEO trained marketers to think in positions: number one, page one, Top 10.

Generative search increasingly requires thinking in distributions.

How often is the brand mentioned across repeated samples? How often is it cited? Which sources recur? How many engines surface it? Does visibility persist next week? Does the same URL survive, or only the domain? How much does the answer change when the model changes?

A single citation cannot answer those questions.

Neither can a single “AI visibility score” unless the methodology behind that score accounts for the underlying instability.

Orbit Media’s 1.7% result is striking because it quantifies how little common ground four major AI systems found in this particular experiment.

But the broader takeaway is more important than the number itself.

AI search is not one new SERP replicated across four brands of chatbot. It is a set of retrieval and generation systems that can look at the same question and construct their evidence from very different corners of the web.

0%