29 AI Overview Citations Shared the Same Content Structure—but Correlation Isn’t a Google Ranking Factor

29 AI Overview Citations Shared the Same Content Structure—but Correlation Isn’t a Google Ranking Factor
Sponsored

A new micro-study offers one of the more useful recent examples of how to think about generative engine optimization — largely because its limitations are almost as important as its findings.

Setproduct examined 29 pages from its own website that had been cited in Google AI Overviews and found a striking set of recurring characteristics. The cited content tended to sit inside deep topical clusters, answer questions in short self-contained paragraphs, use question-form headings, include original data, identify a real author and carry FAQ structured data.

That sounds like a GEO playbook. It is not yet evidence of one.

The Setproduct analysis, published September 8, explicitly acknowledges that it is observational rather than experimental. The publisher did not randomly assign content structures, run controlled A/B tests or isolate the effect of individual variables. It observed patterns across 29 successful pages on one domain and compared them with content that had not received the same citations.

The distinction matters because SEO has a long history of turning correlations into imaginary ranking factors. AI Overviews create the same temptation with a newer vocabulary.

What did the 29 cited pages have in common?

Setproduct reports several structural commonalities across its cited content.

Most citations came from articles within the site’s strongest topical clusters, particularly areas where the publisher had already developed substantial coverage. Every cited article reportedly used question-based H2 headings, contained at least one short direct-answer paragraph early in the content, had FAQ schema and carried a named author linked to a biography page.

The publisher also says every cited article contained at least one first-party data point generated from its own tracking, testing or user base.

Those patterns are interesting because they form a coherent content model rather than a collection of arbitrary SEO tricks: establish subject depth, make questions explicit, answer them clearly, contribute information competitors cannot simply paraphrase and make the source accountable.

What the dataset cannot tell us is which of those characteristics caused the citation — or whether any of them did.

Why doesn’t this study prove that these are Google ranking factors?

The core problem is confounding.

A page with original research may also earn more links. A named expert may publish on a domain with stronger authority. A comprehensive topical cluster may rank better in traditional Search. Pages with good headings may simply be easier for humans to understand. Older articles may have accumulated more impressions, links and crawling history.

If those pages are then cited by AI Overviews, the visible structural feature may correlate with citation while another variable is doing much of the work.

Setproduct itself recognizes this problem. The author says the analysis is not a controlled experiment and states that the site cannot determine the weight Google assigns to any individual signal, whether existing domain authority is a prerequisite or merely an accelerant, or whether the same patterns will persist as AI Overview retrieval evolves.

That is the right level of caution.

Twenty-nine pages from one domain is a useful case study, not a general law

Sample size is another limitation.

Twenty-nine observed citations can reveal hypotheses worth testing, particularly when several pages share the same structure. But all 29 come from the same publisher, meaning they also share domain history, templates, editorial practices, link environment, technical infrastructure and subject-area reputation.

A result observed on one established technology site may not reproduce on a medical publisher, ecommerce store, local business or financial comparison website.

There is also selection bias. The study begins with pages known to have been cited and works backward to identify common characteristics. That is useful exploratory analysis, but it is weaker than prospectively changing one factor across comparable pages and measuring whether citation probability changes.

The appropriate conclusion is therefore “these characteristics appeared repeatedly in this dataset,” not “Google uses these six factors to choose AI Overview citations.”

Deep topical coverage is the strongest strategic hypothesis

Setproduct says the clear majority of its 29 citations came from the domain’s deepest topical clusters. Articles outside those established clusters received far fewer citations even when the individual pages were well optimized.

This observation aligns with a broader content principle that Google itself has documented for years.

Google’s people-first content guidance asks publishers whether a page provides original information, research or analysis and whether it offers a substantial, complete or comprehensive description of its topic. It also encourages sites to demonstrate clear expertise and sourcing.

That does not confirm “topical authority” as a discrete AI Overview ranking factor. It does make comprehensive subject expertise a sensible publishing strategy independently of whether it directly causes an AI citation.

This is an important test for GEO advice: would the recommendation still improve the page if AI Overviews disappeared tomorrow? Deep, useful topical coverage passes that test.

Short direct answers may improve extractability without being a magic word count

Setproduct reports that its cited pages frequently contained direct-answer paragraphs of roughly 40 to 70 words. The author describes these as self-contained answer blocks that can be extracted easily.

The useful part of that observation is structural. A paragraph that answers the question immediately is easier for readers to understand and easier for retrieval systems to interpret than several paragraphs of preamble before the answer appears.

The dangerous part would be converting “40–70 words” into a universal Google specification.

Google has not published such a requirement. Setproduct explicitly describes the range as an industry-observed pattern rather than a Google rule.

Publishers should therefore optimize for answer completeness and clarity, not for a mechanical word counter. Some questions need one sentence. Others require several paragraphs, qualifications or evidence.

Question-based headings are plausible retrieval aids, but exact-match claims remain unproven

Every cited Setproduct article reportedly used question-style H2 headings, and the publisher recommends matching headings closely to the way users phrase queries.

Again, there is a reasonable information-architecture argument behind the recommendation.

A heading such as “How does server-side rendering affect AI crawlers?” tells both the reader and a machine exactly what the following section addresses. A vague heading such as “A technical consideration” communicates much less.

What the study cannot establish is that Google AI Overviews use an exact-match question heading as a special retrieval key or that reproducing a query verbatim creates a citation advantage.

Clear descriptive headings are good editorial structure. Treating exact query syntax as a confirmed AIO mechanism goes beyond the available evidence.

Original data may be valuable because it gives the page something worth citing

One of the most compelling findings is that every cited article reportedly contained at least one first-party data point.

This makes intuitive sense without requiring a secret ranking-factor theory.

If ten pages repeat the same generic advice and one publishes a measured statistic, experiment or original case study, the latter gives a search system a piece of information that cannot be sourced equally well from every competitor.

Google’s own helpful-content guidance explicitly asks whether content provides original information, reporting, research or analysis. Its guidance for succeeding in AI Search similarly emphasizes unique, valuable, non-commodity content.

Setproduct goes further by hypothesizing that AI Overviews prefer specific data partly because it can reduce hallucination risk. That may be plausible, but Google has not documented such a mechanism as a citation-selection rule.

The safer takeaway is simpler: original evidence makes a page more useful and more citable, regardless of the precise retrieval algorithm.

Named authors fit Google’s trust guidance, but author markup is not a citation guarantee

Setproduct also found that every cited article had a named author with a linked biography page. Those author pages included professional information and structured data.

This is another area where the observation overlaps with Google’s public quality guidance.

Google recommends making clear who created content where readers would reasonably expect authorship information. Its people-first guidance specifically points to sourcing, evidence of expertise and background about the author or publishing site as elements that can help users assess trust.

That supports transparent authorship as an editorial practice.

It does not establish that adding Person schema to an author page will cause an AI Overview citation. An identifiable author can correlate with higher editorial standards, stronger expertise, established reputation and better sourcing — all variables that the Setproduct sample cannot separate.

FAQ schema is the easiest finding to overinterpret

Every one of Setproduct’s 29 cited articles reportedly contained FAQ schema. That is visually the cleanest correlation in the study, which also makes it the most tempting to turn into a technical shortcut.

Google’s current documentation argues against doing that.

Search Central says pages do not need special AI-specific markup to appear as supporting links in AI Overviews or AI Mode. According to Google’s AI features documentation, a page needs to be indexed and eligible to appear in Google Search with a snippet; there are no additional technical requirements specifically for AI features.

Google’s structured-data guidelines also state that correct markup creates eligibility for supported Search features but does not guarantee that a feature will appear.

Google has not documented FAQPage schema as an AI Overview citation factor.

That does not mean the correlation is meaningless. FAQ markup may coexist with highly structured question-and-answer content that is itself easy to retrieve. The markup and the content structure are also highly correlated with each other, making it difficult to determine which element matters from this dataset.

Schema correlation is not evidence that AI Overviews are reading an FAQ pipeline

Setproduct proposes that FAQ schema provides AI Overviews with a pre-structured question-and-answer pair that can be extracted without parsing prose.

That is a mechanistic hypothesis, not a Google-published specification.

This distinction matters because marketers frequently convert plausible technical explanations into factual descriptions of systems whose internals they cannot observe.

The evidence in this case is that the 29 cited Setproduct pages had FAQ schema. The evidence is not that Google selected those pages because of the schema, nor that AI Overview retrieval necessarily consumes FAQ markup through a dedicated citation pipeline.

Publishers should implement structured data when it accurately represents visible content and follows Google’s policies, not because a 29-page sample supposedly discovered a hidden ranking switch.

The older-page edits are intriguing but still do not establish causation

Setproduct provides another potentially interesting observation.

The publisher says it revised a small number of older pages that had not been cited, adding FAQ schema and rewriting H2 headings to match query syntax. Some began appearing in AI Overviews in the following weeks, while others did not.

This resembles a before-and-after test, but several variables changed simultaneously and there was no randomized control group.

Time itself is also a variable. The publisher says most of its cited articles did not receive citations during their first few weeks and tended to appear only months after publication, after accumulating organic impressions and backlinks.

Setproduct explicitly says it cannot isolate whether that delay reflects authority accumulation, crawling or other retrieval behavior.

So the post-edit citations are useful evidence for a hypothesis, not proof that the edits triggered them.

Google says AI Overviews use the same foundational SEO principles

The strongest counterweight to emerging GEO checklists is Google’s own documentation.

Google says the foundational practices that apply to ordinary Search also apply to its AI experiences: make pages crawlable and indexable, allow snippets, follow Search policies and create helpful, reliable, people-first content.

It specifically says there are no additional technical requirements for appearing as a supporting link in AI Overviews or AI Mode.

That does not mean AI retrieval behaves identically to ten blue links. Google acknowledges that AI features can use different models and techniques and that the set of responses and links can vary.

It means publishers should be skeptical of claims that one schema type, paragraph length or heading formula is a documented prerequisite for citation.

The best GEO recommendations survive even if the causal theory is wrong

There is a useful way to separate durable optimization from speculative optimization.

Deep topical coverage helps readers understand a subject. Direct answers reduce friction. Descriptive question headings improve navigation. Original research creates differentiated value. Named authors improve accountability. Correct structured data makes content more machine-readable for the Search features that support it.

All of those can be worthwhile even if none turns out to be an independent AI Overview ranking factor.

That is a much stronger foundation for GEO than implementing a tactic solely because a small observational study found it on cited pages.

The more a recommendation depends on an undocumented theory about Google’s internal retrieval pipeline, the more cautiously it should be treated.

A better next experiment would change one variable at a time

Setproduct’s analysis is valuable because it generates testable hypotheses. The next step is stronger experimental design.

A publisher could identify comparable pages within the same topical cluster and randomly change one structural feature on a subset: for example, rewriting headings into direct questions while leaving the rest of the content unchanged.

Another experiment could add accurate FAQ markup without changing the visible text. A third could introduce original first-party data into otherwise comparable articles.

Researchers would then need to track a stable set of queries over time, record whether an AI Overview triggers, whether the page is cited, its citation position or prominence, traditional ranking changes and other confounding variables.

Even that would be difficult because Google’s AI responses can vary between searches and evolve during the experiment. But it would move the evidence closer to causal inference than simply inspecting successful pages after the fact.

The 29-page study is useful precisely because it should not become doctrine

Setproduct has produced a compact dataset with several sensible observations. The cited pages repeatedly show traits that good editors and SEOs would recognize: depth, clarity, specificity, originality and accountable authorship.

Those commonalities are worth testing.

What they do not justify is a new checklist of alleged Google ranking factors.

The publisher itself is unusually clear about this. It calls the work an observational teardown, says it did not run controlled experiments and acknowledges that it cannot determine which signals carry weight or whether existing authority explains part of the result.

That caveat should travel with every screenshot of the findings.

Twenty-nine AI Overview citations sharing a content structure tells us where to look next. It does not tell us why Google selected them.

In GEO, as in SEO, correlation is evidence for a question — not the answer.

0%