In Two GEO Experiments, Third-Party Pages Drove 86% of Brand Mentions—and the Most-Cited Source Wasn’t the One Sending Traffic

In Two GEO Experiments, Third-Party Pages Drove 86% of Brand Mentions—and the Most-Cited Source Wasn’t the One Sending Traffic
Sponsored

Two commercial GEO experiments suggest that brands trying to become visible in AI answers may need to think beyond the content they publish on their own domains. Across two tests that manually tracked 15 commercial-intent queries each and logged 775 citation events, third-party placements repeatedly emerged as important sources of brand visibility—and citation volume did not reliably predict referral traffic or business value.

The experiments, documented by Zeeshan Yaseen in Search Engine Land on September 14, are unusually useful because they followed changes over time rather than relying on a single screenshot. They are also easy to overinterpret. These were two commercial case studies in different niches, without randomized controls, and several interventions were made together. The results reveal associations worth testing, not universal laws of generative engine optimization.

The 85.8% figure comes from the cold-start experiment

The headline number needs a precise denominator. In the second experiment, a SaaS link-building agency began with no measurable presence across the tracked AI platforms. Researchers monitored 15 commercial-intent keywords from May 30 through June 28 across ChatGPT, Claude, Gemini, Perplexity, Google AI Mode and Grok.

The brand accumulated 298 appearances during the 30-day test. When the researchers counted the sources behind those appearances, they logged 437 source mentions. Third-party listicles accounted for 85.8% of those mentions, the brand's own listicle accounted for 14.0%, and PR represented 0.2%.

That does not mean 85.8% of all citations across both experiments came from third-party pages. It is specifically the source mix within the second experiment's 437 mentions. Even with that qualification, the imbalance is striking: earned placements produced far more observed source mentions than the controlled asset on the brand's own site.

A few third-party sources dominated the citation mix

The distribution was highly concentrated. Three sources—Indie Hackers, Bruce Jones SEO and the brand's own listicle—generated 342 of the 437 source mentions, or about 78% of the total. Seven other live placements divided the remainder.

The researchers had built their outreach list by first running the target queries across the AI platforms and recording which domains were already being cited. Five of eight outreach targets later appeared in the final citation mix. Indie Hackers rose from 44 mentions during prospecting to 146 after the placement went live, Bruce Jones SEO rose from 26 to 69, and TechBullion increased from 15 to 37.

But the method did not work consistently. Some targeted sites generated fewer citations after placement, while others had still not appeared by the end of the measurement window. Observing which sources an AI system already uses may be a useful prospecting signal, but it is not a guarantee of future citation.

Adding competitors was followed by a 4-to-49 jump

One of the most provocative observations involved the brand's own listicle. Initially, the page presented the company largely on its own. On June 23, the researchers revised it to include recognized competitors and make the page more comparative.

Mentions of that source subsequently rose from four to 49. The experiment's daily visibility count also increased from 22 on June 23 to a peak of 95 on June 27.

Yaseen interprets this as evidence that the surrounding peer set may matter: an AI system evaluating a commercial category could find a genuinely comparative resource more useful than a self-promotional page. The first experiment produced a directionally similar observation when a listicle lost visibility after recognized experts were removed.

Still, the timing does not establish causality. There was no untreated control page, and the experiments included other PR, guest-post and editorial activity. The safest conclusion is that adding relevant competitors was followed by a large increase in mentions and deserves controlled replication.

The first experiment found rapid citation decay

The earlier experiment tracked another commercial brand across 15 queries on ChatGPT, Claude, Gemini and Perplexity. Its strategy combined listicle placements, PR, guest posts and organic LinkedIn activity.

By the end, the brand appeared for roughly 10 to 12 of the 15 tracked keywords. Search Engine Land reports platform citation counts of 148 for ChatGPT, 96 for Claude, 87 for Gemini and 64 for Perplexity. Listicles represented 72.4% of citations and PR another 24.1%.

But visibility was unstable. Roughly half of the sources stopped being cited within 30 days. One placement fell from 29 mentions to 11 in a week, while another previously consistent listicle declined without direct intervention.

That finding challenges the idea that earning one strong AI citation creates a durable asset. In these tests, citation visibility behaved more like a changing distribution than a permanent ranking position.

The most-cited source was not the one sending traffic

The second experiment also exposed a measurement problem for GEO programs. Indie Hackers generated the highest citation volume, yet its referral traffic remained flat. TechBullion produced fewer citations but increased sessions from one to 64.

In other words, the source that appeared most frequently inside AI answers was not necessarily the source producing the most measurable visits.

This distinction matters because AI visibility dashboards often elevate citation count or share of voice into the primary KPI. Those metrics can describe exposure, but they do not automatically describe user behavior. A source can influence an answer without receiving a click, while a less frequently cited source can generate disproportionately valuable traffic.

One Perplexity visit became a paying customer

The sharpest illustration came from the lowest-volume platform in the second test. Perplexity generated only four brand appearances, yet one referral produced a prospect who later became a paying customer.

During the 30-day window, 18.5% of new users arrived through referral traffic and another 3.25% through GA4's AI Assistant channel. Together those channels represented just over one-fifth of new users in the experiment.

One conversion cannot establish a platform-level conversion rate. It does show why raw visibility rankings can be misleading. A low-volume AI source can still matter commercially if the users arriving from it have strong purchase intent.

Citation volume, referral traffic and revenue are different metrics

The experiments point toward a measurement stack rather than a single GEO score. Citation frequency tells a brand how often a source appears. Brand mentions indicate whether the company enters the answer. Referral analytics show whether users visit. CRM data reveals whether those visits create pipeline or revenue.

Those layers can diverge substantially. A publisher can be highly cited but send little traffic. An AI platform can generate few visits but produce a valuable customer. A brand can become more visible without gaining proportional website sessions.

For teams accountable for revenue, citation count should therefore be treated as an intermediate metric rather than the final outcome.

Owned content looked more like infrastructure than the main growth engine

Yaseen says the second experiment changed his view of owned listicles. After the first test, he had considered publishing comparative listicles on a brand's own domain a stronger primary tactic. The cold-start test suggested a more limited role.

The owned listicle was the slowest source to receive a citation, taking 18 days, while several third-party placements appeared much sooner. It ultimately contributed 14% of the 437 source mentions.

That does not make owned content unimportant. A brand still needs clear, authoritative pages explaining what it is, what it offers and how it compares. But in this case, the larger visibility gains came from being represented on sources outside the brand's direct control.

GEO may be partly a reputation-distribution problem

The common thread across the two experiments is association. The observed AI systems did not appear to evaluate the brand only through its own website. They repeatedly surfaced third-party pages that placed the company inside a category, comparison or expert peer set.

That makes GEO resemble digital reputation management as much as conventional on-page optimization. The question becomes not only, “What does our website say?” but also, “Which sources describe us, who do they compare us with, and are those sources already part of the retrieval ecosystem for our buyers' questions?”

This helps explain why digital PR, editorial placements and comparative content can matter to AI visibility. They create external evidence about how a brand fits into a market.

But the experiments cannot isolate which intervention caused the gains

The methodological limitations are substantial. PR, guest posts, listicle placements, LinkedIn activity and editorial changes overlapped. There were no randomized control groups. The two experiments involved different niches, and the platform sets were not identical: the second added Google AI Mode and Grok.

The queries were manually checked, which is useful for close observation but also encounters the inherent variability of generative systems. Results can change between runs, accounts, locations and model updates.

Because several variables moved together, a rise following a placement does not prove that placement caused the rise. Likewise, the competitor addition and subsequent 4-to-49 increase is a compelling temporal pattern, not a controlled causal estimate.

The findings should therefore generate hypotheses for further experiments rather than become a universal GEO playbook.

The most practical lesson is to start with observed sources

Despite those limitations, one tactic is relatively low-risk: before pitching publications, run the commercial questions that matter to the business and record which sources the relevant AI systems already retrieve or cite.

That produces a prospect list based on observed behavior rather than relying solely on conventional authority metrics. A high-authority domain that never appears for the target query may be less immediately relevant to AI visibility than a specialist source that repeatedly enters the answer set.

The experiments also suggest measuring those sources repeatedly. With roughly half of the first test's sources disappearing from citations within 30 days, a one-time check can create a false impression of durable visibility.

Brands should test comparative usefulness, not manufacture endorsements

The peer-set findings should not be interpreted as permission to stuff famous competitors or experts into pages simply to manipulate AI systems. The stronger editorial interpretation is that useful category content often needs context.

A buyer evaluating an agency, software product or service usually wants alternatives, trade-offs and criteria—not a page insisting that one company is the only answer. A genuinely comparative resource can satisfy that intent better.

If AI systems reward that usefulness, the optimization and the reader's interest are aligned. If competitor names are added without meaningful comparison, the tactic becomes much harder to defend and may not reproduce the observed effect anyway.

Two experiments point to a broader GEO measurement problem

The most valuable result may not be the 85.8% figure itself. It is the disconnect among sources, citations, traffic and customers.

In the cold-start experiment, third-party listicles dominated source mentions. Yet the most-cited source was not the strongest traffic source, and a single visit from a low-volume AI platform became a customer. In the first experiment, many citations disappeared within a month.

Those outcomes make a simple “get more citations” strategy look incomplete. GEO performance has to be evaluated over time and connected to actual business outcomes.

For now, these cases are evidence, not proof. Two uncontrolled commercial experiments cannot establish that third-party pages will drive 86% of AI visibility for every brand, that adding competitors causes a 12-fold citation increase, or that any particular platform will convert better.

What they do show is that AI visibility can be distributed across an ecosystem the brand does not own—and that the source most visible to the model may not be the source most valuable to the business. That is a strong reason for GEO teams to measure mentions, citations, referrals and revenue separately before deciding what is actually working.

0%