Ranking in Google Search and being cited by Google’s generative search features may be far less closely connected than many publishers assume. A large empirical study of 11,500 queries found that the sources retrieved by traditional Google Search, AI Overviews and Gemini 2.5 Flash overlapped surprisingly little, suggesting that visibility in one Google surface does not automatically translate into visibility in another.
The paper, “How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews”, was submitted to arXiv on April 30, 2026 and accepted to ACM SIGIR 2026. Researchers Riley Grossman, Songjiang Liu, Michael K. Chen, Mike Smith, Cristian Borcea and Yi Chen compared the sources returned for the same queries across Google’s conventional search results, AI Overviews and Gemini Flash 2.5. Their central finding is not simply that generative search rearranges familiar results. In many cases, it appears to draw from a substantially different source set.
Only 18% overlap between Search and AI Overviews
The researchers measured source similarity using the Jaccard index, which compares the intersection of two sets with their combined unique elements. Across the benchmark, the average similarity between Google Search and AI Overviews was 0.18. The overlap between Google Search and Gemini was 0.16, while AI Overviews and Gemini were even farther apart at 0.11.
Those figures matter because they challenge a simple model of AI search in which a generative answer merely summarizes the pages already winning the traditional SERP. The study instead indicates that the retrieval pipelines can expose substantially different groups of publishers. A page that performs strongly in organic search therefore cannot be assumed to have the same probability of appearing as an AI citation, and success in Gemini does not necessarily imply visibility in AI Overviews either.
The researchers also found systematic differences in the kinds of domains retrieved. Traditional Google Search was significantly more likely to surface popular websites and institutional sources from government or education, while the generative systems were more likely to retrieve Google-owned content. That distinction adds another layer to the emerging question of how publishers should evaluate visibility as search becomes a collection of related but distinct retrieval experiences.
AI Overviews appeared on more than half of representative real-user queries
The study is also notable for the frequency with which AI Overviews appeared. Among 5,000 queries drawn from a dataset intended to represent real user searches, an AI Overview was generated for 51.5%. Across the full 11,500-query benchmark, the paper reports an even higher overall generation rate, although prevalence varied considerably by query type.
AI Overviews were especially associated with longer informational questions, while the researchers found much lower prevalence for some trending queries. This reinforces an important distinction for search analysis: an average AI Overview rate across a large query set does not mean every vertical, intent or keyword portfolio experiences the feature at the same frequency.
Small query changes can produce bigger changes in AI results
Source diversity was not the only difference. The researchers report that AI Overviews were less consistent across repeated runs of the same query and less robust to small changes in query syntax than traditional search. In practical terms, two semantically similar formulations may lead to more variation in generative retrieval than publishers are accustomed to seeing in conventional rankings.
This has consequences for how AI visibility is measured. Traditional rank tracking already requires care because results vary by location, device, personalization and time. Generative search adds another source of variability: the retrieval and answer-generation process itself may be less stable. A single prompt test or one captured AI Overview is therefore weak evidence for broad conclusions about a site’s visibility.
What the study means for SEO and GEO
For publishers and search marketers, the most important implication is that conventional rankings and generative citations should increasingly be measured as separate outcomes. Traditional SEO remains relevant because Google Search continues to distribute traffic and because generative systems still depend on retrievable web content. But an 18% average source overlap between Search and AI Overviews indicates that organic ranking data alone may be an incomplete proxy for AI visibility.
The findings also complicate the developing field of generative engine optimization, or GEO. If source selection differs across Search, AI Overviews and Gemini, there may be no single optimization tactic that reliably improves visibility across every surface. Measurement may need to become query-set based and multi-system: tracking whether a domain is retrieved, cited and represented accurately across repeated prompts rather than treating one AI citation as equivalent to a stable search ranking.
The paper additionally reports that sites blocking Google’s AI crawler were significantly less likely to be retrieved by generative systems. The authors highlight this in the context of publisher choices over AI access and the broader economic relationship between content creators and generative search providers. It is an area where technical crawling controls, retrieval behavior and publisher strategy increasingly intersect.
A crucial limitation: the data are from December 2025
The numbers should not be read as a snapshot of Google’s systems today. The benchmark data were collected on December 7 and 8, 2025, and the generative product examined was Gemini Flash 2.5. Google changes its search and AI systems rapidly, so source selection, AI Overview prevalence and stability may have shifted since the experiment.
That limitation does not make the study less useful. Its strongest contribution is structural rather than predictive: at the time of measurement, three closely related Google discovery surfaces produced markedly different source sets for identical queries. The public 11,500-query benchmark also gives future researchers a basis for testing whether those gaps narrow, widen or change as generative search evolves.
For site owners, the takeaway is not to abandon traditional SEO in favor of an entirely separate AI playbook. It is to stop assuming that one visibility metric represents every search experience. If generative search continues to select sources differently from the classic SERP, publishers will need to monitor organic rankings, AI citations and assistant visibility as connected but distinct channels—and test them repeatedly rather than relying on a single query or snapshot.