Nearly 30% of Google AI Overview Sources Didn’t Rank on Page One—and 11% of Claims Lacked Citation Support

Nearly 30% of Google AI Overview Sources Didn’t Rank on Page One—and 11% of Claims Lacked Citation Support
Sponsored

Google AI Overviews do not simply summarize the websites visible at the top of the same search results page. A large-scale measurement study found that 29.8% of domains cited by AI Overviews did not appear among the co-displayed first-page organic results for the same query, suggesting that Google's generative answer layer draws from a source pool or prioritization process that is meaningfully different from the conventional ranking users see below it.

The study, submitted to arXiv on May 13, 2026, analyzed 55,393 trending Google queries across 19 topical categories during a 40-day period from March 13 through April 21. Researchers Haofei Xu, Umar Iqbal and Jacob M. Montgomery captured AI Overviews, their embedded references, the first-page search results displayed alongside them and the contents of cited webpages.

The sourcing result was only half of the story. After decomposing the generated answers into 98,020 atomic factual claims, the researchers classified 11.0% as inconsistent with the content they could retrieve from the pages Google cited. Most of that failure was omission rather than direct contradiction: 7.0% of claims were not addressed in the retrieved source text, while 2.7% were classified as incorrect and 1.4% as ambiguous.

At the same time, the study found that AI Overview citations came from domains that scored as more credible on average than the first-page results displayed alongside them. That creates a more complicated picture than either “Google cites the best-ranking pages” or “AI Overviews use worse sources.” The system appears able to reach beyond page one for relatively credible domains, yet selecting credible domains does not guarantee that every generated claim is actually supported by those pages.

The researchers measured 55,393 live Google queries over 40 days

The paper, “Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact,” is a longitudinal measurement study rather than a fixed benchmark. Every 24 hours, the researchers collected trending U.S. queries from Google Trends across 19 categories including health, politics, science, technology, shopping, finance, entertainment and travel.

They then issued those queries to Google using fresh, stateless Chrome browser sessions through a Puppeteer-based crawler running on AWS Lambda in Northern Virginia. The crawler detected whether an AI Overview appeared, expanded the generated response, recorded its reference URLs and captured the first-page search results shown for the same query.

Across the full dataset, Google displayed AI Overviews for 13.7% of queries. Activation varied sharply by query form: question-style searches triggered an AI Overview 64.7% of the time, compared with 9.5% for non-question queries. The paper also reports substantial topical variation, reinforcing that AIO exposure cannot be inferred from a single aggregate trigger rate.

This query design is relevant when interpreting the results. The dataset was built from trending searches, not a random sample of every Google query. It therefore captures a broad and naturalistic set of topics receiving real-world attention during the observation period, but it should not be treated as a perfect representation of the global Google query distribution.

Nearly 30% of cited domains were absent from page one

The most important result for SEO and GEO is the limited overlap between conventional rankings and AI citations. The researchers report that 29.8% of domains cited by AI Overviews did not appear in the co-displayed first-page results for the same query.

That does not mean those pages were completely unranked in Google. A cited source could have appeared on page two or deeper, could have ranked for a related query, or could have been surfaced through a retrieval system that is not exposed directly in the visible SERP. The experiment establishes only that almost three in ten cited domains were absent from the first page the crawler observed alongside the AIO.

The finding nevertheless weakens a common assumption that AI Overview visibility is simply conventional top-10 SEO with a generative interface added on top. If nearly 30% of cited domains are not present on page one, then ranking there is neither a necessary condition for being cited nor a complete map of the sources available to the generative layer.

The authors interpret this as evidence that AI Overviews use a source-selection mechanism distinct from Google's visible ranking algorithm. The wording should be kept precise: the study observes different output sets; it does not have access to Google's internal retrieval architecture and cannot identify every mechanism responsible for that difference.

Page-one ranking and AI citation are related visibility surfaces, not identical ones

For practitioners, this distinction changes how AI visibility should be measured. Traditional rank tracking answers where a page appears in the organic SERP. Citation tracking answers whether Google's generative system selected that page as evidence. The study shows those outcomes overlap substantially but not completely.

A site can therefore gain AI Overview exposure without holding a first-page organic position for the exact query. Conversely, ranking on page one does not guarantee selection as an AIO source. The generative system has another opportunity to choose among or beyond the pages that conventional search surfaces to users.

This is one reason GEO cannot be reduced to checking whether existing top-ranking pages are cited. A useful audit needs both datasets: the organic result set and the generative citation set. The gap between them can reveal sources that Google's AI layer values even when conventional rankings do not make that preference obvious.

For competitive research, those out-of-page-one citations may be particularly informative. They identify domains whose content is entering the answer layer despite weaker visible placement, offering clues about topical authority, specificity, freshness or other characteristics that merit closer qualitative analysis.

The cited domains were more credible on average

The researchers also compared source credibility using a continuous domain-quality score derived from expert and crowd assessments of news-domain credibility. The measure covers 11,520 domains and produces scores from zero to one.

Under that metric, domains selected by AI Overviews were systematically more credible on average than the first-page results displayed alongside them. This finding differs from some earlier research that used other source-quality taxonomies and found weaker AIO source quality.

The discrepancy is a reminder that “quality” is not a directly observable universal variable. Different studies use different domain lists, scoring systems and query populations. The paper's result should therefore be stated in relation to its chosen credibility measure rather than converted into a general claim that every AI Overview source is better than every organic result.

Even within the study, higher average domain credibility did not eliminate factual grounding problems. That separation between source reputation and claim support becomes one of the paper's most important findings.

The researchers broke AI answers into 98,020 atomic claims

To test whether citations actually supported what AI Overviews said, the authors did not score each answer as a single unit. They decomposed the generated text into atomic, self-contained factual claims, splitting compound statements and resolving pronouns so each assertion could be checked independently against its cited source.

The final verification dataset contained 98,020 claims from 7,491 verifiable AI Overviews. Ninety-two AIOs were excluded from this part of the analysis: 80 had no extractable claims and 12 relied only on social-media references that the researchers' crawler did not attempt to retrieve.

Each claim was compared with the text collected from the associated cited webpage. The verification framework used Grok 4.1 Fast Reasoning at temperature zero, supported by task-specific few-shot prompts, and the researchers manually validated samples of both the claim-extraction and verification pipeline.

Using an LLM as part of the evaluator introduces its own measurement uncertainty, so the reported 11% should not be mistaken for a perfect ground-truth hallucination rate. The value is the output of a defined and validated automated classification framework applied at unusually large scale.

Eleven percent of claims were not adequately supported by their citations

The verification system classified 84.6% of claims as clearly supported and another 4.4% as vague but consistent, producing an overall consistent share of 89.0%. The remaining 11.0% were classified as inconsistent with the retrieved cited content.

That inconsistent group contains different failure modes. The paper reports 2.7% of all claims as incorrect, meaning the source contradicted or conflicted with the generated assertion. Another 1.4% were ambiguous. The largest category, 7.0%, was omission: the claim was not addressed in the source text available to the verification pipeline.

This breakdown matters because “11% unsupported” can otherwise sound as though one in nine claims directly contradicted its citation. That is not what the researchers found. Direct conflict represented a smaller share; missing support was the dominant problem.

For users, however, omission is still meaningful. A citation displayed beside an assertion visually implies evidentiary support. If the cited page does not contain the claimed information, the user cannot verify the statement from the source Google presented, regardless of whether the statement might be true elsewhere.

Some apparent omissions may reflect inaccessible source content

The authors explicitly document limitations in their webpage collection. Short-form social platforms including Reddit, YouTube, LinkedIn, TikTok, Facebook, Instagram, Pinterest, Threads and X were excluded from content retrieval because authentication and bot protections made them difficult to crawl reliably.

Fewer than 1% of collected pages failed text extraction, while 262 pages contained paywalls. For paywalled sources, the researchers retained only the visible material, typically the headline and opening paragraphs. They acknowledge that a claim supported deeper behind a paywall could consequently be classified as unsupported by the automated checker.

The crawler also used readability extraction to isolate primary page content and remove navigation and boilerplate. As with any automated extraction system, information embedded in unusual page structures could be lost. These limitations mean the 7.0% omission category is best interpreted as claims unsupported by the source text the pipeline could retrieve, not proof that every cited page lacked the information in every possible rendering.

The 2.7% incorrect category is conceptually stronger because it represents conflict with the retrieved source rather than simple absence. Both categories are relevant, but they carry different levels of evidentiary certainty.

Better sources did not guarantee better grounding

One of the study's more consequential conclusions is that source credibility and claim fidelity were largely independent. Selecting a reputable domain did not ensure that the generated sentence accurately reflected what that domain said.

This exposes two separate quality problems in AI search. Source selection asks whether the system retrieves trustworthy evidence. Grounding asks whether the generator faithfully uses that evidence. Improving the first cannot automatically solve the second.

For Google, that distinction is particularly important because the visible citation interface encourages users to infer a connection between a claim and an authoritative source. A reputable publisher's logo can increase confidence even when the particular cited page does not substantiate the sentence beside it.

For publishers, it creates a reputational issue as well as a traffic issue. A publication can be presented as evidence for a statement it did not actually make, potentially lending its credibility to an unsupported synthesis.

The study also found more than half of cited pages carried display advertising

The researchers extended their analysis beyond sourcing and factual support to publisher economics. At least 50.6% of AI Overview-cited pages contained display advertising under the study's detection method.

This matters because AI Overviews can use information from an ad-supported page while satisfying the user's query before a visit occurs. In that scenario, the publisher contributes content to the answer but loses the pageview that would have generated an advertising impression.

The paper contrasts that displacement with Google's own sponsored search inventory, which can continue appearing on pages containing AI Overviews. The authors frame this as an economic asymmetry: generative answers can reduce the organic referral opportunity available to publishers while Google's advertising remains inside the search experience.

The study does not directly calculate revenue lost per citation or establish that every AIO impression prevents a click. Its advertising analysis identifies economic exposure—the proportion of cited pages whose business model can depend on traffic—rather than a precise dollar estimate of harm.

Question-style searches were dramatically more likely to trigger AI Overviews

The 64.7% activation rate for question-form queries provides additional context for content teams. Users who formulate searches as explicit questions were 6.8 times more likely to encounter an AI Overview than users issuing non-question queries in the study.

That means informational content answering natural-language questions is particularly exposed to the generative layer. The pages most carefully designed to answer “how,” “why,” “what” and similar questions may face the highest probability that Google synthesizes their information before the user reaches an organic result.

It also makes question-based prompt monitoring more relevant for GEO. A site can hold stable conventional rankings while its actual user journey changes because a growing share of natural-language queries receives a generated answer above those results.

Activation was not uniform across topical categories, however, and the paper reports lower rates for politically sensitive topics. Marketers should therefore avoid applying the 13.7% overall rate or the 64.7% question rate mechanically to every vertical.

The study measures one 40-day period in a rapidly changing product

All observations were collected between March 13 and April 21, 2026. That temporal boundary is important because Google continually changes AI Overview triggering, retrieval, models, citation interfaces and supported markets.

The results are a large-scale measurement of the system operating during that window, not a permanent specification of AI Overviews. A future replication could find different activation rates, source overlap or claim-support performance after product changes.

The crawler also used U.S.-localized Google results from Northern Virginia without personal browsing history. Real users may see differences based on location, language, personalization, device or experimental treatment. The controlled setup improves consistency while limiting claims about every possible Google experience.

For GEO practitioners, that volatility is itself a lesson. Citation studies should be dated and methods should be recorded. A percentage without a collection window can become misleading quickly in an AI search product that changes continuously.

Ranking reports alone cannot explain AI Overview visibility

The 29.8% source-overlap result has a direct operational implication. If an SEO team monitors only first-page rankings, it will miss a meaningful portion of the domains feeding Google's generated answers.

That does not make rankings obsolete. More than 70% of cited domains did overlap with page-one results under the study's comparison, showing substantial continuity between traditional and generative visibility. Conventional search relevance remains an important part of the source ecosystem.

But the non-overlap is large enough that citation monitoring needs to become a separate analytical layer. Teams should record which URLs and domains AI Overviews cite, compare them with organic positions and investigate the sources that consistently enter the generative layer without top-page placement.

Those sources may reveal opportunities that a rank-only competitive analysis misses. They may also reveal that Google is selecting a different page from the same domain than the one optimized for conventional search.

GEO needs to measure support quality, not just citation counts

The 11% inconsistency finding points to another measurement problem. Being cited is usually treated as a positive outcome in GEO dashboards, but citation frequency alone says nothing about whether the brand or publisher is represented accurately.

A source can be cited for an unsupported claim. It can be cited next to an ambiguous synthesis. Its reporting can be compressed in a way that changes the original meaning. For organizations with reputational or regulatory exposure, those outcomes may be more important than the raw number of citations.

AI visibility programs should therefore sample the generated claims associated with important citations and verify them against the source page. This is especially important for medical, financial, legal and product claims where a subtle mismatch can carry consequences beyond lost traffic.

The ideal GEO metric is not simply “share of citations.” It includes whether the citation is prominent, whether the claim is faithful, whether the brand is represented correctly and whether users have a viable path to the original evidence.

Google AI Overviews are creating a second editorial layer above ranking

Traditional search gives Google editorial power through ranking: it decides which pages users see first, but users still choose which source to open and interpret. AI Overviews add another layer. Google selects evidence, synthesizes it into prose and places that synthesis above the conventional results.

The study shows that this new layer is not merely a textual rendering of page one. Nearly 30% of cited domains were absent from the first-page results shown for the same queries, and the cited domains were, on average under the study's measure, more credible than the surrounding result set.

Yet the generation stage introduced its own fidelity gap. Of 98,020 atomic claims, 11.0% could not be classified as adequately supported by the retrieved cited content. Better source selection and faithful source use are therefore separate engineering challenges.

For SEO and publishers, the implication is equally two-sided. Ranking still matters, but it no longer fully predicts who supplies the answer. Earning a citation can create visibility beyond page one, but the citation may not produce a click and may not even guarantee that the resulting claim faithfully represents the source. In Google's AI search era, the optimization target is expanding from where a page ranks to whether it is selected, how it is interpreted and whether users ever reach the evidence behind the answer.

0%