Google Search, AI Overviews and Gemini Often Choose Almost Completely Different Sources for the Same Query

Google Search, AI Overviews and Gemini Often Choose Almost Completely Different Sources for the Same Query
Sponsored

Google Search, AI Overviews and Gemini may all belong to the same company, but they frequently construct their answers from strikingly different parts of the web. A large comparative study accepted at SIGIR 2026 found average source overlap of only 18% between traditional Google results and AI Overviews, 16% between Google results and Gemini, and just 11% between AI Overviews and Gemini.

The study, submitted to arXiv on April 30, introduces a public benchmark of 11,500 queries and compares the sources returned by Google's traditional search engine, the AI Overview displayed with the SERP and Gemini 2.5 Flash. The researchers also examine how often AI Overviews appear, how stable each system is when the same or slightly modified query is repeated, which types of domains receive visibility and whether blocking Google's AI crawler is associated with reduced inclusion.

The result challenges one of the most convenient assumptions in generative engine optimization: that strong Google rankings provide a reliable map of the sources Google's AI products will use. Traditional ranking clearly remains relevant, but the overlap figures show that each surface creates a substantially different source set. Even AI Overviews and Gemini—the two generative systems in the comparison—shared less source overlap with each other than either did with the conventional SERP.

There is an equally important limitation. The researchers collected results on December 7 and 8, 2025, using SerpAPI for Google results and Gemini 2.5 Flash for the assistant comparison. Google changes Search, AI Overviews and Gemini continuously. The paper is therefore a rigorous snapshot of those product versions and collection methods, not a permanent specification of how Google's systems retrieve sources today.

The benchmark compares 11,500 queries across three Google surfaces

The authors—Riley Grossman, Songjiang Liu, Michael K. Chen, Mike Smith, Cristian Borcea and Yi Chen—built the benchmark from several query subsets designed to capture different search behaviors. The collection includes 5,000 queries from ORCAS, a dataset derived from real user search behavior, alongside question-answering, shopping, debate and explanation-oriented query sets.

For every query, the researchers collected the traditional Google search results, any AI Overview returned with the SERP and a response from Gemini 2.5 Flash. They extracted the sources associated with each system and compared the resulting lists using Jaccard similarity, which measures set overlap, and rank-biased overlap, which also accounts for the position of results.

The distinction between those metrics is useful. Jaccard similarity asks how much two source sets share regardless of order. Rank-biased overlap asks whether the systems agree particularly on the sources placed near the top, where user attention is usually concentrated. The study reports the same broad pattern under both approaches: the systems retrieve substantially different source lists.

Search and AI Overviews shared only 18% average source overlap

The average Jaccard similarity between traditional Google Search and AI Overviews was 0.18. In intuitive terms, only about 18% of the combined unique sources returned by those two surfaces for a query appeared in both sets.

This does not mean that 82% of AI Overview citations were necessarily absent from Google's first page; Jaccard similarity is calculated over the union of the two source sets and is not equivalent to a one-directional citation percentage. The correct interpretation is that the two lists, taken together, had relatively little overlap.

That distinction matters when comparing this paper with other AIO citation studies. A study asking “what share of AIO citations also rank on page one?” uses a different denominator from a study asking “how similar are these two sets overall?” Both can demonstrate divergence between generative and conventional search, but their percentages are not interchangeable.

For SEO teams, the practical conclusion survives the statistical nuance. A top Google result set is an incomplete proxy for the sources an AI Overview may select. Ranking analysis and citation analysis need to be tracked separately.

Gemini diverged from traditional Search even more

The average overlap between traditional Google results and Gemini was 0.16, slightly below the 0.18 measured for Search and AI Overviews. Gemini therefore drew from a source set that was also substantially different from the conventional SERP.

This matters because “Google AI visibility” is often discussed as though it were one measurable outcome. The paper shows at least three distinct surfaces with different retrieval behavior: organic Google Search, AI Overviews embedded in Search and Gemini as an assistant experience.

A brand can perform strongly in one and weakly in another without contradiction. The systems may use related infrastructure and belong to the same company while applying different retrieval, ranking, grounding and presentation logic.

That makes cross-product GEO reporting more complicated. A dashboard that measures Gemini citations cannot automatically be treated as a measurement of AI Overview visibility, just as a conventional rank tracker cannot fully represent either generative surface.

AI Overviews and Gemini had the lowest overlap of all

The most surprising comparison is between Google's two generative experiences. AI Overviews and Gemini shared an average Jaccard similarity of just 0.11, the lowest pairwise overlap in the study.

The paper notes the result as notable because AI Overviews themselves use Gemini-family technology. But sharing model technology does not require sharing an identical retrieval stack. The products operate in different contexts, can have different tool access and may optimize for different answer formats, latency constraints or product goals.

For practitioners, the 11% figure is a warning against optimizing for an abstract category called “LLMs.” A source pattern observed in Gemini may tell little about which pages AI Overviews will cite for the same query. Even within one vendor, generative visibility is product-specific.

This also limits the value of one-off citation screenshots. If two Google generative products can return largely different source sets for identical questions, a single result from one surface cannot stand in for a brand's broader AI visibility.

AI Overviews appeared on 51.5% of representative real-user queries

The study also measured how frequently AI Overviews appeared. Across the 5,000-query ORCAS subset, which the authors describe as most representative of real user search behavior, Google generated an AI Overview for 51.5% of queries.

Across the complete 11,500-query benchmark, the prevalence was higher at 65.6%. The difference illustrates how heavily AIO trigger rates depend on query composition. A benchmark containing many informational, explanatory or question-style searches will produce a different aggregate rate from one dominated by navigational or transactional behavior.

The researchers found that informational intent, question formatting and query length were associated with higher AIO prevalence. Questions triggered AI Overviews much more often than keyword-style searches in the benchmark, while longer queries were also more likely to receive generated summaries.

That makes the 51.5% ORCAS result more useful for broad behavioral context than the 65.6% full-benchmark figure, but neither should be interpreted as Google's permanent global AIO coverage rate. Both belong to the December 2025 collection window and the study's query distribution.

Generative search favored Google-owned properties more often

The paper also compares which kinds of domains each system retrieves. Traditional Search was significantly more likely to surface popular websites and institutional government or educational sources, while the generative systems were more likely to retrieve Google-owned content.

YouTube is the most obvious example of a Google property that can benefit from that shift. The broader finding matters because generative search does not merely reorder the same external websites. It can alter the composition of the information ecosystem presented to users.

The authors frame this as an important question for competition and publisher visibility. If AI interfaces increasingly answer queries directly while also favoring first-party properties in their source selection, external publishers face two pressures at once: fewer direct navigational opportunities and a changed probability of being selected as evidence.

The study does not establish why each individual Google-owned source was selected. The observed domain preference could reflect content format, availability, retrieval design, product integration or other factors. It demonstrates a systematic difference in the output, not the internal causal mechanism.

Popular traditional-search domains often lost visibility in generative results

The source-composition analysis also found that some domains with strong traditional search presence appeared less frequently in the generative systems. The paper highlights examples including Reddit, Wikipedia and Amazon, while some long-tail or less dominant websites gained relative visibility.

That shift has an important GEO implication. Generative retrieval can create new opportunities for sites that do not dominate conventional rankings, because the system may value a different set of source characteristics when grounding an answer.

But opportunity is not the same as predictability. The low overlap and higher variability measured elsewhere in the study mean those gains can be less stable than a conventional rank position. A site may enter a generated answer without acquiring a durable, easily monitored ranking equivalent.

For competitive analysis, marketers should therefore look beyond the domains that dominate organic page one. The generative source set can reveal niche publishers, specialist sites and first-party properties that are nearly invisible in a conventional SERP comparison.

Sites blocking Google's AI crawler appeared less often in AI Overviews

One of the paper's more sensitive findings concerns Google-Extended, the robots.txt token publishers can use to control certain uses of their content by Google's Gemini ecosystem. The researchers observed that websites blocking Google's AI crawler were significantly less likely to be retrieved by AI Overviews than by traditional Google Search.

The result is notable because the paper argues that AI Overviews technically have access to content indexed through traditional Google crawling. In other words, the reduced AIO presence of blocking sites was not explained simply by those pages disappearing from Google Search.

The authors also identify a group of popular publishers that appeared repeatedly in Search and AI Overviews but were never cited by Gemini; the publishers in that group blocked Google-Extended. That pattern is consistent with the documented purpose of Google-Extended for Gemini grounding.

Care is required before turning the AIO finding into a causal SEO rule. The study is observational at the domain level. Sites that block AI crawling may differ systematically from sites that allow it, and the paper does not demonstrate that adding or removing the directive alone causes a specific AIO visibility change.

The result nevertheless makes crawler policy a legitimate variable to include in GEO audits. Publishers deciding whether to block AI use should recognize that access controls can interact with visibility differently across Search, AI Overviews and Gemini.

AI Overviews were less consistent across identical runs

Source divergence was not limited to comparisons between products. The researchers also tested internal consistency by running the same queries multiple times. AI Overviews produced less consistent source sets than traditional Google Search.

This is a familiar property of generative systems but a significant change for search measurement. Traditional rankings fluctuate, yet the same query repeated within a short window usually retains a substantial core of results. Generative retrieval can introduce another layer of variability in which sources appear even when the user's wording does not change.

For GEO tools, this means one execution is a weak measurement unit. A brand cited once may disappear on the next run without any underlying change to the page. Reliable reporting needs repeated samples and a frequency-based metric rather than a binary cited/not-cited flag.

The study's consistency results also help explain why marketers can disagree about whether a particular domain is “ranking” in AI. They may be observing different valid generations of a system whose source selection is inherently less deterministic.

Small query edits caused larger changes in AIO sources

The researchers further tested robustness by making minor changes to query syntax. AI Overviews were less robust than traditional Search: small modifications could produce larger changes in the retrieved source set.

That finding makes prompt formulation an important measurement variable. Two queries that a human considers semantically equivalent can expose different source ecosystems in a generative system. A GEO audit built around one exact wording can therefore miss substantial parts of the visibility landscape.

Traditional SEO already accounts for keyword variants, but generative search raises the dimensionality. Word order, question framing, added context and conversational phrasing can affect not only ranking position but whether the answer is generated and which sources the system retrieves.

A practical prompt set should therefore contain semantically related variants and repeat them over time. The goal is to estimate a distribution of visibility rather than declare one deterministic rank.

GEO needs product-specific metrics

The study makes a strong case against a universal “AI rank.” With average source overlap of 18% between Search and AIO, 16% between Search and Gemini and 11% between AIO and Gemini, each surface exposes a different version of the web.

A useful reporting system should separate them. Organic Search needs conventional rankings and click data. AI Overviews need trigger frequency, citation share and referral behavior. Gemini needs its own prompt-level citation and brand-representation tracking. Combining those outcomes into one score can hide the very differences the paper documents.

Cross-surface analysis is still valuable, but it should answer a different question: how portable is a site's visibility? A domain appearing consistently across Search, AIO and Gemini may possess source characteristics that survive different retrieval systems. A domain visible on only one surface may be more dependent on that product's current design.

The findings also complicate generative engine optimization experiments

GEO experiments often modify a page and then measure whether it appears more frequently in one model's answers. The paper shows why that result cannot automatically be generalized to other Google surfaces.

An optimization that improves Gemini visibility might have no effect on AI Overviews because the products retrieve different sources. A change that preserves organic ranking might still alter generative retrieval. A crawler policy could affect Gemini differently from conventional Search.

The correct experimental unit is therefore the product and query family, not “Google” or “AI” in the abstract. Teams should document the model version, date, location, query wording and collection method whenever they report a GEO gain.

This is particularly important because the study itself is now historical in product terms. Gemini 2.5 Flash and the December 2025 versions of AI Overviews are not frozen systems. Reproducing the benchmark today could yield different absolute overlap values even if the broader pattern of divergence remains.

The December 2025 snapshot is a major limitation—and a methodological lesson

All primary results were collected over two days, December 7 and 8, 2025. The researchers used SerpAPI to collect Google Search and AI Overview results and the Gemini API for Gemini 2.5 Flash. That gives the study a clearly defined and reproducible collection window, but it also sharply bounds what the numbers mean.

Google has had months to modify AI Overview retrieval, source presentation, triggering and model infrastructure since those observations. Gemini has also evolved. The 18%, 16% and 11% overlap figures should therefore be cited with the date and model rather than presented as timeless properties of Google's current products.

This limitation is common across AI-search research and should become a reporting standard. Every citation study needs a timestamp. Without one, readers cannot distinguish a stable architectural pattern from a temporary product configuration.

SerpAPI is another methodological boundary. It provides a scalable way to collect results, but an API-mediated snapshot is not identical to every user's personalized browser experience. Location, device, account state and live Google experiments can affect what an individual sees.

The bigger lesson is that generative search creates multiple visibility markets

For decades, marketers could treat the Google SERP as the dominant map of search visibility. Different result features existed, but the ranked organic list remained the central reference point for which websites Google considered relevant.

AI Overviews and Gemini fragment that map. The same query can now produce three source populations under one corporate umbrella, with surprisingly little overlap. Traditional Search favors one set of institutions and popular domains; AI Overviews construct another evidence set; Gemini can draw from a third.

That fragmentation creates opportunity for sites that struggled to break into conventional page one, but it also makes visibility harder to predict and measure. A strong organic position no longer guarantees generative inclusion, while a generative citation may be unstable across runs or disappear after a minor query edit.

The SIGIR study's most useful contribution is therefore not a new target percentage for GEO. It is evidence that “Google visibility” has become a plural concept. In the December 2025 snapshot, Search and AI Overviews shared only 18% average source overlap, Search and Gemini 16%, and the two generative products just 11%. For publishers and brands, the same query can lead Google Search, AI Overviews and Gemini to almost completely different parts of the web. Any serious search strategy now has to measure those surfaces separately before it can understand how they interact.

0%