252,000 GEO Tests Found Four Citation Gatekeepers—and Formatting Wasn’t One of Them

252,000 GEO Tests Found Four Citation Gatekeepers—and Formatting Wasn’t One of Them
Sponsored

Generative Engine Optimization has accumulated a long list of tactical advice: add headings, restructure paragraphs, bold important facts, write in a more quotable style and hope an AI system prefers the page. A large controlled experiment presented at SIGIR 2026 suggests that some of the strongest citation effects may be much less cosmetic.

In “What Gets Cited: Competitive GEO in AI Answer Engines”, Rahul Vishwakarma, Shushant Kumar and Ratnesh Jamidar ran 252,000 paired trials across six large language models. Their test changed one content factor at a time and measured which of two competing documents received the first citation in the generated answer. Four factors emerged as unusually consistent across the models: topical relevance, the source’s position in the supplied context, explicit price information and a recent timestamp. Formatting-only changes, by comparison, produced little impact.

The scale and controlled design make the paper unusually useful for GEO research, but its biggest limitation is just as important as its headline result. The researchers did not test live Google AI Mode, ChatGPT Search, Perplexity or another production answer engine. They supplied two documents directly to each model in a simulated retrieval-augmented generation environment, meaning the experiment begins after the retrieval problem has already been solved.

The experiment isolates citation competition rather than search retrieval

The researchers wanted to answer a narrow question: if an AI system already has two candidate sources available, what makes it cite one before the other? Real-world observational studies struggle to isolate that effect because competing pages differ simultaneously in authority, relevance, age, structure, brand recognition, backlinks and dozens of other characteristics. A citation difference in a live engine therefore rarely proves which characteristic caused the outcome.

The SIGIR experiment removes much of that ambiguity. Each trial supplied exactly two candidate documents to the model, with the pair differing in one tested factor. Brands and publishers were anonymized, and source order was counterbalanced to help separate content effects from position bias. Across 18 factors and six LLMs, the researchers recorded which source appeared in the first citation marker and analyzed the resulting preferences with mixed-effects models.

The six tested models were Gemini 2.5 Flash, GPT-5 Nano, GPT-5 Mini, GPT-5.2, Claude 3.5 Sonnet and Kimi K2 Thinking. That variety is useful because it allows the study to distinguish effects that appear in one model from patterns that persist across different model families.

Topical relevance is the clearest content gatekeeper

The strongest content lesson is also the least surprising: the source needs to be about the thing the user is asking about. The paper identifies topical relevance as one of the biggest drivers of first-citation selection across all six models. When an on-topic source competed against an off-topic alternative, the relevant document received a major advantage.

That finding sounds obvious, but it matters because GEO advice can become preoccupied with page-level tricks after relevance has been taken for granted. The experiment suggests that cosmetic optimization cannot compensate for a source that is fundamentally a poor match for the question. Citation selection begins with usefulness to the requested topic, not with whether the publisher has applied a particular visual template.

For practitioners, the important interpretation is not “repeat the keyword more often.” Topical relevance in an answer-engine context means that the information supplied to the model actually addresses the task it is trying to complete. A page can be semantically related to a category while still failing to contain the facts needed for a specific recommendation or comparison.

Source position produces a huge effect—but it is easy to misread

The second major finding is position. Models strongly preferred the source presented first in the supplied list or tool context. This is a striking result because the underlying document did not need to become better; simply changing where the source appeared could change which one was cited first.

However, “position” in this paper does not mean ranking number one on Google. It means the ordering of the two documents that the experimental system placed into the LLM’s context. The researchers deliberately counterbalanced that order precisely because language models can exhibit positional bias.

This distinction is central to any GEO interpretation. In a production AI search engine, another retrieval or ranking system may determine which documents enter the context and in what order. The SIGIR study demonstrates that context ordering can influence the downstream citation decision once documents are present; it does not reveal how a publisher earns that position in Google, ChatGPT, Perplexity or another live system.

Explicit price information consistently helps in the tested scenarios

Price was another factor that moved citation selection consistently across the six models. When one candidate included explicit pricing and the competing source did not, the document containing the price was more likely to receive the first citation.

This result fits the functional role of citations in product-oriented answers. If a user is evaluating products and the model needs to support a statement about cost, a source that contains a concrete price gives it evidence that the alternative cannot provide. The value comes from supplying a decision-relevant fact, not from a mysterious preference for currency symbols.

The domain matters, however. The experiment uses product-oriented scenarios, so publishers should not generalize the result into a rule that every page needs a price. A medical explanation, historical article or software troubleshooting guide may have no legitimate pricing information to provide. The broader lesson is that specific facts required to answer the query can create citation utility.

Recent timestamps also win consistently

Freshness produced another cross-model effect. Documents carrying a recent timestamp were consistently favored over old alternatives in the relevant comparison. For information that can change—prices, product availability, recommendations, specifications or market conditions—a newer source gives the model a rational reason to prefer it.

This finding should not be converted into a “change the date” growth hack. The controlled experiment measures date variants under specific conditions; it does not prove that republishing stale information with a 2026 timestamp will improve citations in live AI engines. Production systems can use other freshness signals, and users ultimately need information that is current in substance rather than merely recent in metadata.

The defensible GEO takeaway is that temporal clarity matters when recency is relevant to the task. Publishers should make genuinely current information identifiable as current and avoid forcing models to guess whether a time-sensitive fact is still valid.

Completeness and trust cues help, but less consistently

Below the four most consistent factors, the study found smaller gains from characteristics such as completeness and trust cues. These effects were not as uniform or dominant across models, but they reinforce a familiar pattern: a source that provides enough evidence to answer the question and gives the model reasons to treat that evidence seriously can outperform a weaker alternative.

The model-to-model variation is itself informative. There may not be one universal citation preference shared at identical strength by every LLM. Different models can react differently to coverage depth, social proof, specifications, evidence and other tested attributes, which makes simplistic “optimize once for every AI engine” checklists difficult to defend.

The paper’s precise odds ratios also vary dramatically by model, and some statistical fits carry convergence warnings. The most robust reading is therefore the directional hierarchy across the experiment rather than treating any single effect size as a permanent law of AI citations.

Formatting-only optimization produced little return

One of the most useful negative findings concerns formatting. The researchers tested formatting-oriented changes and found little consistent effect on which source received the first citation. In the controlled setup, changing presentation without materially changing the information did not compete with the effects of relevance, position, price or recency.

This is a useful corrective to a growing GEO cottage industry built around cosmetic page transformations. Bolding sentences, changing layout or making text look more “AI friendly” is easy to sell because it creates a visible before-and-after change. The study suggests those edits may have limited value once the model already has access to the full documents and is deciding which one to cite.

But this result has a boundary too. The experiment bypassed live crawling and retrieval. Formatting and structure can still affect accessibility, parsing, extraction and human usability before a document reaches an LLM context. The paper therefore does not prove that site structure is irrelevant to production AI search; it shows that the tested formatting-only differences had little influence in the post-retrieval citation competition the researchers measured.

The biggest real-world variable was deliberately removed

The study’s clean causal design depends on eliminating the messy first half of AI search. The model does not crawl the web, decide which domains to trust, retrieve ten candidates and then narrow them down. The experiment hands it the complete text of two candidate sources and asks it to answer from that controlled context.

That means the study cannot tell a publisher how to get retrieved in the first place. Domain authority, crawlability, indexing, search-engine rankings, link graphs, retrieval embeddings, platform partnerships and other upstream systems may determine whether a page ever reaches the citation competition tested here.

This is not a flaw so much as the reason the experiment can isolate content factors cleanly. The danger comes only when the results are presented as a complete model of live AI visibility. In production, a perfectly optimized document that never enters the retrieval set has no opportunity to benefit from any downstream citation preference.

GEO has at least two optimization layers

The paper therefore supports a useful way to divide GEO work. The first layer is retrieval eligibility and competitiveness: can the engine discover, access, understand and select the source for the context window? The second is citation competitiveness: once the source is present alongside alternatives, does it contain the information that makes the model want to reference it?

The SIGIR experiment is unusually strong evidence about the second layer. Relevance, decision-useful facts such as explicit prices, clear freshness and the source’s position in the supplied context all influenced first-citation selection. Cosmetic formatting did comparatively little.

The first layer remains much harder to study because every production engine has its own retrieval architecture and those systems can change without notice. Publishers should resist using post-retrieval evidence as if it were proof of retrieval-ranking factors.

First citation is a narrow but meaningful metric

The researchers measure which source receives the first citation marker. That is a deliberately simple outcome that enables hundreds of thousands of controlled comparisons, but it is not the only visibility outcome that matters in an AI answer.

A source cited second may still receive meaningful exposure. A page can influence an answer without being cited visibly. A brand can be recommended even when a third-party source receives the citation. And citation does not automatically translate into clicks, revenue or customer preference.

For GEO teams, first-citation probability is therefore best treated as one measurable component of a larger visibility system. The study advances our understanding of citation choice without claiming that first citation equals business impact.

The research is stronger than a live-engine correlation study—and narrower

Observational AI visibility studies can analyze thousands or millions of live citations, which makes them valuable for understanding what production engines actually surface. Their weakness is confounding: high-ranking pages may simultaneously be newer, more authoritative, more relevant and better linked, making it difficult to determine why the engine chose them.

This study takes the opposite trade-off. By changing one factor at a time, it can make much stronger statements about the behavior of the tested LLMs under controlled conditions. In exchange, it sacrifices the complexity of the real retrieval pipeline.

Neither method makes the other obsolete. Live-engine studies tell marketers what is happening in the wild; controlled experiments can help explain which variables cause changes when other conditions are held constant. GEO research needs both.

The practical lesson is to improve information before decoration

For publishers, the most defensible action from 252,000 trials is not a new formatting template. It is a prioritization principle.

First, make the page genuinely relevant to the question it is meant to answer. Then provide concrete information the model needs to support its response, including explicit commercial facts where appropriate. Keep time-sensitive information substantively current and make that freshness clear. Build complete, trustworthy resources rather than assuming visual tweaks will manufacture citation preference.

After that, test formatting and structure for accessibility, retrieval and usability where they make sense—but do not confuse an easy cosmetic edit with evidence-backed GEO leverage.

The paper’s most important contribution may be that it makes the boundary of its evidence visible. Across six LLMs and 252,000 controlled trials, relevance, context position, explicit price and recency repeatedly changed which source was cited first, while formatting alone did little. That is strong evidence about what happens after two sources reach an LLM. The harder problem for every publisher remains what happens before that moment: getting into the retrieval set at all.

0%