AI search engines are designed to make web research easier by synthesizing information and attaching citations, but a new audit raises a more complicated question: what happens when the sources behind those answers were themselves generated by AI? A preprint submitted to arXiv on May 22, 2026 found evidence of AI-generated material among citations returned by ChatGPT, Copilot, Gemini and Perplexity, with approximately 16% of the successfully analyzed unique cited sources classified as AI-generated.
The study, “Synthetic Sources?: Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources,” was conducted by Mowafak Allaham and Nicholas Diakopoulos of Northwestern University. The researchers audited the four generative search systems using 712 English-language, human-generated queries covering politics, health and the environment—areas where source quality can have consequences well beyond routine product discovery or entertainment searches.
What the researchers actually measured
The 712-query benchmark included 175 politics queries, 257 health queries and 280 environment queries. Each query was submitted through the user-facing interface of ChatGPT with web search enabled, Copilot, Gemini and Perplexity. Rather than evaluating only the generated answers, the researchers collected the sources the systems cited, producing a dataset of 26,266 unique URLs across 7,675 domains.
Not every URL could be analyzed. The researchers successfully extracted the main text from 19,154 sources, or 72.9% of the unique URLs in the citation dataset. Pages that could not be processed included non-text formats such as PDFs and images, inaccessible pages and content that had been removed. That distinction is important: the headline figure is not a direct measurement of every citation produced in the experiment.
To identify synthetic text, the authors evaluated AI-detection tools and ultimately used Pangram. Among the 19,154 successfully scraped sources, 2,916 were classified as “Highly Likely AI” and another 140 as “Likely AI.” Combining those two categories produced 3,056 sources, approximately 16% of the analyzed set. The researchers excluded the more ambiguous “Possibly AI” category from the main estimate in an effort to avoid inflating the result.
All four AI search engines cited synthetic sources
The audit found evidence of AI-generated sources across every provider tested, but the proportions differed substantially. Among sources that could be scraped and classified, the paper reports 27.8% for Copilot, 14.7% for Gemini, 9.4% for Perplexity and 7.3% for ChatGPT. These provider-level numbers should be interpreted within the experiment rather than as permanent rankings of the products: generative search systems, retrieval pipelines and web indexes can change quickly.
The distribution of citations was also highly uneven. Across the four systems, the 25 most frequently cited domains accounted for 23.8% of citations, while much of the remaining activity was spread across a long tail of sites that appeared only once or twice. Wikipedia was the most cited individual domain in the top group, while government websites collectively represented a substantial share of the most frequently cited sources.
The synthetic material was especially concentrated outside that small group of dominant domains. According to the paper, only 2.9% of the sources classified as AI-generated came from the top 25 domains, while 97.1% came from the rest of the domain distribution. That creates a difficult quality-control problem: answer engines can offer broad source diversity, but evaluating thousands of rarely cited sites for provenance and reliability is much harder than maintaining a compact whitelist of established publishers.
The risk is not simply that a source used AI
AI-generated text is not automatically inaccurate, and the authors explicitly acknowledge that synthetic content is not necessarily low quality. The concern is instead about provenance and compounding uncertainty. A generative search engine may retrieve a webpage written partly or entirely by another model, synthesize that material into a new answer and then present the original page as supporting evidence. If the underlying synthetic page contains an error, weak sourcing or a hallucinated claim, citation can give that material an appearance of external verification it may not deserve.
This dynamic is particularly relevant to health, politics and environmental information. Users often treat citations as a signal that a generated statement can be traced to independent evidence. But when the cited page is itself machine-generated, the chain of evidence may be more circular than it appears. The challenge for AI search is therefore not merely to attach links, but to assess where those links came from, how their content was produced and whether they ultimately rest on authoritative evidence.
Source concentration and the long tail matter for GEO
The findings also have implications for publishers studying visibility in generative engines. The audit suggests that AI citation ecosystems combine two apparently conflicting patterns: a relatively narrow set of domains receives citations repeatedly, while a very large long tail of domains receives only occasional exposure. In the aggregate data, 59.1% of cited domains appeared only once and another 16.5% appeared twice.
For generative engine optimization, or GEO, that means citation visibility may be less analogous to holding a stable organic ranking than it first appears. A domain can enter an answer engine’s source pool without becoming a consistently preferred authority. Conversely, repeatedly cited domains may gain disproportionate influence over the information synthesized for users. Measuring AI visibility therefore requires more than recording whether a site appeared once; frequency, query coverage, topic and source context all matter.
Important limitations of the 16% figure
The paper is a preprint, and its result should be read with several qualifications. The 712 queries focus on three topics and were filtered for U.S. relevance, so they do not represent every type of question people ask AI systems. The environment queries also came from a different underlying dataset than the politics and health queries. The audit examined user interfaces rather than APIs, which improves relevance to the consumer experience but also ties the observations to the specific product versions and behavior available during testing.
The classification of sources also depends on an automated AI-text detector rather than definitive authorship records. The researchers benchmarked Pangram and GPTZero on their own samples of human-written and newly generated texts before choosing a detector, but AI-content detection remains an inferential method. In addition, only 72.9% of unique cited URLs could be scraped for classification. The roughly 16% figure therefore describes the successfully analyzed portion of this particular citation dataset, not a universal rate for AI search.
Even with those caveats, the study highlights an emerging problem for the web’s information supply chain. Generative engines increasingly act as intermediaries between users and publishers, while the web they retrieve from contains a growing volume of machine-produced material. If AI systems begin repeatedly summarizing, citing and redistributing content created by other AI systems, the distinction between primary evidence and synthetic interpretation becomes harder for users to see.
The practical lesson is that citations alone are not enough to establish trust. For users, publishers and developers, the next stage of AI search quality will depend not only on whether an answer provides sources, but on whether those sources have credible provenance and connect back to reliable evidence. This audit does not establish how often today’s versions of ChatGPT, Copilot, Gemini or Perplexity cite synthetic material across the entire web. It does show that, in the researchers’ sample, AI-generated sources were present across all four—and often hidden behind the reassuring appearance of a citation.