Restructuring Content Lifted AI Citation Rates 17.3% Across Six Generative Engines

Restructuring Content Lifted AI Citation Rates 17.3% Across Six Generative Engines
Sponsored

What if improving visibility in AI search required changing less of what an article says and more of how it is built? A new research preprint argues that document architecture, information chunking and visual emphasis can materially affect whether generative engines cite a source, even when the underlying meaning is preserved.

The study introduces GEO-SFE, short for Structural Feature Engineering for Generative Engine Optimization. In experiments spanning 200 articles and six generative engines, the authors report an overall 17.3% improvement in citation performance after structural optimization, alongside an 18.5% average improvement in a model-based subjective evaluation. The paper was submitted to arXiv on March 31, 2026 and should be read as experimental research rather than a set of established ranking factors.

That caveat is essential. The reported percentages come from the researchers' specific dataset, transformations, evaluation design and tested systems. They do not mean that changing heading depth, paragraph length or bold formatting on a production website will automatically produce a 17.3% citation increase. What the paper does offer is a structured hypothesis for a question publishers increasingly care about: can the presentation of information influence how AI search systems retrieve, parse and cite it?

GEO-SFE separates structure from meaning

Much of the early discussion around Generative Engine Optimization has focused on semantic changes: adding statistics, quotations, authoritative language or clearer factual statements. GEO-SFE takes a different route. Its objective is to alter structural features while preserving semantic integrity, allowing the researchers to investigate whether organization itself affects citation behavior.

The framework divides content structure into three levels. Macro-structure describes the architecture of the document, including heading hierarchy, navigation, logical progression and cross-references. Meso-structure covers section-level organization such as paragraphs, lists, tables, information density and chunking. Micro-structure addresses sentence-level presentation, including emphasis markers, keyword positioning and syntactic patterns.

This distinction is useful because generative search does not necessarily consume a webpage the way a human reader does. A person can skim a long page, interpret visual hierarchy and mentally connect distant passages. Retrieval and generation systems may instead break documents into chunks, rerank fragments, summarize sections or repeatedly search for additional evidence. The structure surrounding a fact can therefore affect whether the system encounters it in a usable context.

Document architecture produced the largest share of the reported gain

The paper's ablation analysis attempts to isolate how much each structural layer contributes. In the authors' experiments, macro-structure accounted for 44.9% of the total improvement, meso-structure for 39.7% and micro-structure for 15.4%. That makes document architecture and information organization considerably more important in the reported setup than visual emphasis alone.

At the macro level, GEO-SFE measures features such as heading hierarchy, consistency among sections, logical progression and internal cross-references. The underlying idea is that a document with a coherent hierarchy gives retrieval and generation systems stronger cues about how concepts relate to one another. Cross-references can also support systems that perform iterative or multi-hop retrieval rather than treating each page fragment independently.

The finding pushes against an overly simplistic version of GEO in which publishers search for a few formatting tricks. In this experiment, the largest effect was associated with the architecture that organizes the entire document, not merely with bolding important phrases. If the result generalizes, information architecture may be a more durable optimization target than cosmetic markup.

Chunking matters because AI systems often consume fragments

The meso layer focuses on how information is divided inside sections. GEO-SFE measures paragraph organization, lists, tables, format diversity and information density. The researchers then transform content toward structural configurations their models predict will improve citation probability while enforcing semantic-preservation constraints.

This aligns with a practical reality of retrieval-augmented generation: a generative engine may not feed an entire 2,500-word article into the final answer-generation step. Search and retrieval pipelines frequently operate on passages or chunks. A useful fact buried in a large, multi-topic block may therefore be harder to isolate than the same fact presented inside a coherent passage whose surrounding sentences provide enough context to stand on their own.

The authors propose specific numerical ranges for paragraph length and structured-format usage, but these should be treated with particular caution. The paper, for example, describes 150–300 words as an optimized paragraph-length range within its framework and suggests a target proportion for lists, tables and other structured elements. Those are experimental parameters derived from the study's methodology, not universal specifications for publishers. Real systems differ in retrieval methods, chunk sizes, interfaces, models and continual product updates.

Visual emphasis is a signal, not a magic citation button

At the micro level, the framework studies emphasis markers and the position of important terms within sentences and structural boundaries. The researchers assign different weights to formatting such as bold and italic text and attempt to place emphasis around important information while maintaining readability.

Micro-structure contributed the smallest portion of the overall gain in the ablation study, at 15.4%. That is still meaningful within the experiment, but it also provides a useful warning against reducing GEO to formatting hacks. The research does not show that randomly bolding keywords causes AI engines to cite a page. Rather, emphasis operates as one component of a larger structural system that includes document hierarchy and information chunking.

There is also an obvious human constraint. A page engineered so aggressively for machine parsing that every other sentence is highlighted, boxed or converted into a list may become unpleasant to read. GEO-SFE explicitly includes semantic and readability constraints, reinforcing a principle publishers should preserve regardless of whether the framework's citation gains reproduce elsewhere: structural optimization should improve information access without corrupting meaning or usability.

The study tested 200 articles across six systems

The experimental dataset contains 200 articles sampled from GEO-bench across six domains: biography, health, technology, finance, travel and science. The articles were paired with 377 real-world queries and averaged 2,547 words. The authors evaluated original and structurally optimized versions, producing 2,400 test cases across six platforms.

The paper groups the tested systems into three architectural categories. Search-then-synthesize is represented by Google SGE and Bing Chat; iterative-refinement systems by Perplexity AI and Phind; and integrated search-generation by ChatGPT and Claude. The authors argue that different retrieval-generation architectures respond to different structural characteristics, so the framework includes architecture-specific weighting rather than assuming every generative engine behaves identically.

That architectural distinction may be one of the study's most important ideas. A system that retrieves documents once and then synthesizes an answer can value structural signals differently from a system that repeatedly searches and refines its evidence. Likewise, a real-time retrieval process may favor passages that can function more independently. GEO may therefore be less like optimizing for one stable algorithm and more like designing content that remains intelligible across several retrieval architectures.

The reported citation rate rose from 45.0% to 52.8%

The headline 17.3% figure is a relative improvement reported across the experiment. In the paper's aggregate results, baseline citation rate was 45.0%, while GEO-SFE optimized content reached 52.8%. Search-then-synthesize systems moved from 43.7% to 52.1%, iterative-refinement systems from 52.3% to 59.6%, and integrated search-generation systems from 39.1% to 46.8%.

The authors report statistical significance at p<0.001 and a Cohen's d of 0.64 for the aggregated citation results. Within their test environment, the improvements were therefore not presented as isolated anecdotal wins. They appeared across all three architectural groups, although the magnitude varied.

Still, a controlled benchmark is not the open web. Commercial AI products change frequently, may personalize responses, use different indexes, apply undisclosed source-quality systems and vary retrieval behavior by query. Repeating the same experiment months later—or against another set of topics and websites—could produce different numbers. The 17.3% figure is evidence from this study, not a guaranteed uplift available to any page that adopts the framework.

What does the 18.5% subjective-quality gain actually mean?

The second headline number needs even more careful interpretation. The authors report an 18.5% average improvement across seven subjective dimensions including relevance, influence, uniqueness, position, citation count, click probability and diversity. However, this was not a conventional panel of human readers rating the pages. The paper says the subjective evaluation used G-Eval with Gemini 2.5 Pro, taking median scores from repeated samples.

That does not make the result meaningless, but it changes what can be claimed. It is better described as a model-based subjective assessment than proof that human users found the optimized pages 18.5% better. The largest reported gains in that evaluation were for influence and estimated click probability, while diversity improved by a smaller amount.

For publishers, this distinction is another reason not to convert experimental metrics directly into business forecasts. Citation rate, model-assessed quality, referral clicks, conversions and reader satisfaction are related but different outcomes. A structural change that helps an AI system extract a passage is valuable only if it also fits the publisher's editorial, commercial and user-experience goals.

What publishers can reasonably test

The paper provides a stronger case for testing structural clarity than for copying its numerical thresholds. Publishers can audit whether long pages have a logical heading hierarchy, whether each section answers a recognizable sub-question, whether important facts are buried inside mixed-topic paragraphs, and whether tables or lists genuinely make structured information easier to extract. They can also examine whether internal references help readers and retrieval systems understand relationships among concepts.

Those changes have an advantage over many speculative GEO tactics: they can improve human readability at the same time. Clear hierarchy, coherent chunks and well-chosen structured elements are useful even if a particular AI citation algorithm changes. That makes structural quality a more defensible investment than chasing an undocumented preference of one generative engine.

Measurement should also remain platform-specific. A publisher testing structural changes should track whether pages are actually cited more often for a defined query set, whether the citations are accurate, whether AI referral traffic changes and whether downstream engagement improves. Without that measurement loop, “AI-friendly structure” can quickly become another collection of unverified best practices.

A research result, not a new universal GEO rulebook

GEO-SFE offers a useful contribution because it treats content structure as an experimental variable rather than an editorial intuition. Its three-level model—document architecture, information chunking and visual emphasis—provides a vocabulary for investigating why two semantically similar documents might perform differently in generative search.

But the most tempting details in the paper are also the ones that should be generalized least aggressively. Specific heading depths, paragraph ranges, formatting ratios and emphasis percentages come from the authors' framework and experimental environment. They should not be presented as confirmed ranking factors for ChatGPT, Google, Perplexity, Claude or any other commercial system.

The defensible conclusion is narrower and more interesting: in this preprint's evaluation, restructuring content while attempting to preserve its meaning increased aggregate citation performance by 17.3% across six tested generative engines, and document-level architecture plus information chunking accounted for most of the reported effect. If future independent work reproduces that pattern, GEO may increasingly look less like adding special phrases for AI and more like an old editorial discipline viewed through a new technical lens: organize information so clearly that both people and machines can find the right piece at the right moment.

0%