Generative search has made citation counts one of the easiest new GEO metrics to track. A brand or publisher can count how many sources ChatGPT, Google or Perplexity displays and treat a higher number as greater visibility. A new arXiv study argues that this can miss a second, fundamentally different outcome: a source can be cited without contributing very much to the generated answer, while another source may shape multiple parts of the response.
The paper, “From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms”, analyzes a public dataset containing 602 controlled prompts, 21,143 valid search-layer citations, 23,745 citation-level feature records and 18,151 successfully fetched pages across ChatGPT, Google AI Overview/Gemini and Perplexity. Its central distinction is between citation selection—the decision to surface a source—and citation absorption, the degree to which observable language, evidence or structure from that page appears to be reflected in the generated answer.
The cross-platform pattern is striking. Perplexity averaged 16.35 citations per prompt and Google 12.06, while ChatGPT averaged only 6.88. Yet among pages successfully fetched for the second stage of analysis, ChatGPT’s mean influence score was 0.2713, compared with 0.0584 for Google and 0.0646 for Perplexity. In other words, the platform with the narrowest citation breadth showed by far the highest value on the paper’s proxy for source-level contribution. That does not prove ChatGPT literally “reads more deeply” internally, but it does show why counting visible links and measuring apparent answer contribution should not be treated as the same metric.
“Absorption” is an observable proxy, not a window into the model
The qualification matters because the researchers do not have access to hidden attention states, private retrieval traces or causal records showing which document generated which token. Their influence score is constructed from observable characteristics of the final answer and cited page. It combines repeated references to a source, how early the citation appears, the share of answer paragraphs associated with it, TF-IDF similarity and bigram/trigram overlap. A page that repeatedly aligns with the language and coverage of an answer therefore receives a higher score than a source that appears once with little textual relationship to the response.
That makes the metric useful for comparing citation behavior, but it also defines its limits. The score cannot establish that the model would have produced a different answer if the source had been removed, and it is not a direct measurement of internal model attention or retrieval importance. The authors explicitly describe the study as observational and separate their descriptive findings from the confirmatory statistical models that would be needed for stronger inference. The dataset also is not a probability sample of all real-world AI searches, and the absorption analysis covers only successfully fetched pages; the reported overall fetch-success rate was 76.44%.
Those constraints make “citation ≠ influence” a better interpretation than “ChatGPT uses every source more deeply.” The latter turns a proxy into an internal-mechanism claim the data cannot support. What the study can show is that, under its prompt design and measurement formula, ChatGPT answers exhibited substantially more observable source-answer overlap per fetched citation even while presenting fewer citations overall.
Definitions, numbers and procedures are associated with deeper answer participation
The content patterns provide a second useful result. Pages with higher influence scores tended to be longer, more modular and more semantically aligned with the generated response. They were also more likely to contain extractable evidence genres such as definitions, numerical facts, comparisons and procedural steps. These are information units that can be reused naturally inside a synthesized answer: a model can define a concept, insert a statistic, contrast two options or reproduce the logic of a process while still writing a new response around that evidence.
The authors describe this as an “evidence-container” hypothesis. In that framing, GEO has two gates. A source first needs enough authority, recognizability, language fit and domain context to enter citation selection; once selected, it needs semantic alignment, structural legibility and dense usable evidence to contribute substantially to the answer. The distinction helps explain why a highly authoritative page might be cited as background while a more precisely structured page supplies the facts that dominate the response.
It is still a hypothesis derived from descriptive associations rather than a controlled content experiment. The paper does not demonstrate that adding a definition or comparison to an existing page will cause its influence score to rise on a live platform. Longer and more structured pages may differ in many other ways, and semantic alignment is naturally entangled with relevance. The practical value is therefore in identifying characteristics worth testing, not converting them into guaranteed GEO ranking factors.
The Q&A result challenges one of GEO’s easiest shortcuts
One of the paper’s most useful negative findings concerns Q&A formatting. If answer engines respond to questions, it is tempting to assume that converting pages into explicit question-and-answer blocks automatically makes them easier to absorb. The observed data does not support that shortcut. Q&A pages recorded a mean influence score of 0.0947, compared with 0.1005 for non-Q&A pages—a 5.74% relative difference in the opposite direction.
That result does not mean FAQs are harmful or that question-based headings should be removed. It means the wrapper is not the evidence. A short Q&A block containing a generic answer may be less useful to a generative system than a conventionally structured section containing a precise definition, quantified comparison or complete procedure. The finding reinforces the paper’s broader distinction between formatting for apparent machine readability and supplying information that can actually support a synthesized response.
For publishers, this changes how GEO performance can be reported. Citation share measures whether a source enters the visible evidence set; an absorption-style metric asks how extensively the answer appears to draw on it. A page can succeed at one and fail at the other. That creates at least two separate optimization questions: why is the source not being selected, and, once selected, why is it contributing little to the answer?
Citation breadth and citation depth should be separate dashboard metrics
The platform comparison makes the reporting implication especially clear. Perplexity’s 16.35 citations per prompt suggests a broad-source strategy in this dataset, while ChatGPT’s 6.88 suggests a narrower one. Simply ranking the platforms by citation count would therefore make Perplexity look dramatically more source-rich. But the influence proxy reverses the picture: ChatGPT’s 0.2713 mean is more than four times the corresponding values reported for Google and Perplexity. The two measures are describing different behaviors rather than contradicting each other.
That distinction can also prevent poor editorial decisions. If a publisher optimizes only for citation frequency, it may pursue pages and domains that are repeatedly listed but barely reflected in answers. If it optimizes only for textual absorption, it may overlook the harder first-stage problem of getting selected at all. A useful GEO dashboard should preserve both dimensions, then connect them to downstream outcomes such as referral traffic, branded demand, leads or conversions rather than compressing every stage into one visibility score.
The study does not settle how AI engines choose or use sources, and its influence metric should not be mistaken for direct access to a model’s reasoning. Its contribution is more foundational: it gives researchers and marketers a vocabulary for distinguishing a link that is displayed from a source that appears to shape the answer around that link. In generative search, those are no longer safely interchangeable outcomes.
That makes the headline finding less about which platform is “better” at citations and more about what GEO should measure next. Perplexity and Google may expose a wider source set, while ChatGPT may show deeper observable overlap with each source it does cite. For publishers, the strategic question is therefore no longer just “Did the AI cite us?” It is also “What, if anything, from our page made it into the answer?”