OpenAI Licensing Deals Are Linked to 48% More ChatGPT Citations

OpenAI Licensing Deals Are Linked to 48% More ChatGPT Citations
Sponsored

A new dataset is putting a number on a question publishers have been asking since AI companies began signing content agreements with major media groups: do those commercial relationships change which sources AI systems cite?

A joint study from Press Ranger and OtterlyAI matched 129.3 million AI citations from more than 20 million URLs against 91 confirmed licensing agreements between AI companies and publishers. Its headline finding is striking. Publishers with an OpenAI licensing agreement averaged 10.2 ChatGPT citations per cited page, compared with 6.9 for publishers without one — a reported 48% difference. Among publishers that had signed exclusively with OpenAI, the gap reached 112%.

Those numbers deserve attention. They do not, however, justify the simplistic conclusion that publishers can pay OpenAI and buy citation placement. The study is observational, not a controlled experiment, and the publishers selected for licensing deals are unlikely to be a random sample of the web. The more interesting question for GEO is therefore not whether licensing “buys citations.” It is whether AI source selection is driven purely by relevance at retrieval time, or whether privileged content-access relationships can become another variable in what an AI system sees, trusts and ultimately exposes to users.

What the 129.3 million-citation study found

Press Ranger and OtterlyAI say they examined citation activity across seven AI platforms during June 2026: ChatGPT, Google AI Overviews, Google AI Mode, Perplexity, Microsoft Copilot, Gemini and Claude. Press Ranger supplied a database of publicly confirmed licensing agreements, while OtterlyAI supplied the citation dataset.

The ChatGPT result was the clearest platform-specific difference. OpenAI-licensed publishers received an average of 10.2 citations per cited page, versus 6.9 for publishers without a deal. Across all seven platforms combined, the same OpenAI-licensed group averaged 10.7 citations per cited page compared with 7.3 for unlicensed publishers, a reported 46% gap.

The exclusive-OpenAI cohort produced an even more provocative result. Publishers that had signed with OpenAI but not with another AI company recorded 112% more ChatGPT citations per cited page than publishers without licensing agreements. According to the study, 57.9% of the AI citation volume received by OpenAI-licensed publishers came from ChatGPT itself.

That concentration is what makes the result more interesting than a generic correlation between large publishers and high AI visibility.

The effect does not appear equally across other licensing platforms

If licensing agreements simply identified publishers that were already larger, more authoritative and more citation-worthy everywhere, we might expect similar advantages across the platforms signing those deals. The reported results are more uneven.

In the study's comparison, publishers with Google licensing agreements did not show a corresponding citation premium in Google AI Overviews; they were reportedly cited slightly less often than comparable unlicensed publishers. Perplexity-licensed publishers were approximately at parity with unlicensed publishers on Perplexity. The pronounced “home-platform” association therefore appears strongest in the OpenAI/ChatGPT pairing.

That asymmetry does not prove preferential treatment. It does make the data harder to dismiss as nothing more than a universal publisher-authority effect. Something about the OpenAI-licensed cohort, its content, its access relationship with OpenAI, or the way ChatGPT retrieves and integrates those sources appears different enough to warrant testing.

Licensing can change access without changing a ranking algorithm

There is an important mechanism that sits between “pure organic relevance” and “paid citation placement”: content access.

OpenAI's own publisher announcements make clear that at least some partnerships are designed to bring licensed journalism directly into ChatGPT. In its 2026 announcement with Grupo Folha and Grupo UOL, for example, OpenAI said the partnership would bring the publishers' journalism to ChatGPT, allowing users to see summaries based on their reporting with attribution and links to original sources. OpenAI described the arrangement as part of a broader effort to integrate trusted reporting into AI-powered experiences.

That distinction matters. A licensing agreement does not need to contain a secret “boost this publisher” instruction to affect visibility. Cleaner access, structured feeds, contractual permission to use content, deeper archive availability, more reliable attribution metadata or product integrations could all alter the effective candidate set from which an AI system generates answers.

If one source is consistently and reliably available to a system while another must be discovered through ordinary web retrieval, the two sources may not be competing under identical conditions even if the final citation-selection logic is relevance-based.

This is AI preference versus AI visibility

For NetContentSEO, this suggests a useful distinction. AI visibility measures whether a publisher or brand appears. AI preference asks whether a system selects that publisher more often than its observable relevance, authority and content supply would predict.

The difference is crucial. A large publisher should naturally earn more AI citations if it produces more useful pages, covers more topics and has stronger authority. That is visibility explained by supply and relevance. Preference begins when two publishers with comparable eligible content systematically receive different treatment inside one AI platform.

The Press Ranger/OtterlyAI data does not establish that second condition because it does not fully control the variables needed to estimate an expected citation rate. But it gives us a candidate signal. The next step should be to test whether the 48% gap remains after accounting for publisher size, topic coverage, content type, freshness, authority and the number of pages actually eligible to answer the same prompts.

The metric itself needs careful interpretation

There is another methodological detail worth emphasizing: the headline comparison is citations per cited page. That is not the same as asking what percentage of all published pages receive citations.

A publisher could have a relatively small subset of pages entering the citation dataset but see those pages cited repeatedly. Another could have a broader range of pages cited less frequently each. Without the complete denominator of eligible or published content, we should be careful about translating 10.2 versus 6.9 citations per cited page into a general probability that any given article will be selected.

The research is also vendor-produced. Press Ranger sells PR and distribution services oriented toward search and AI visibility, while OtterlyAI sells AI search monitoring. That does not invalidate the dataset, but it increases the importance of transparent methodology and independent replication. A finding this commercially and strategically significant should eventually be tested by researchers with access to comparable citation-scale data.

The biggest confounder: OpenAI chooses whom to license

Licensing deals are not randomly assigned. OpenAI has generally partnered with established media organizations whose content already has attributes likely to correlate with citation: editorial authority, large archives, frequent publication, recognizable brands and coverage of high-demand topics.

This creates a classic selection problem. Perhaps OpenAI-licensed publishers are cited more because the agreement improves access. Perhaps they are cited more because OpenAI chose publishers ChatGPT was already likely to cite. Perhaps both effects operate simultaneously.

The 112% figure for OpenAI-exclusive publishers does not eliminate that problem. Exclusivity may define a particular class of publishers with its own characteristics. Without before-and-after citation data around deal dates and a matched control group, causality remains unresolved.

That is why “OpenAI deals are linked to more citations” is the defensible headline. “OpenAI deals cause more citations” is not yet supported.

A stronger experiment: what happens before and after a licensing deal?

The cleanest next test would use time. For every publisher with a known agreement date, measure ChatGPT citation frequency before and after the partnership and compare that change with a matched group of publishers that did not sign a deal. Matching should account for publication volume, topical mix, domain authority, geography and baseline citation rate.

If citations rise materially after the deal while comparable publishers remain stable, the access hypothesis becomes stronger. If the citation advantage existed at roughly the same level before the contract, selection bias becomes a more plausible explanation: OpenAI may simply be licensing publishers its systems already use heavily.

A second experiment could compare the same publisher across AI platforms. If an OpenAI partner receives a post-deal citation lift primarily on ChatGPT but not on Gemini, Perplexity or Claude, that would provide stronger evidence of a platform-specific relationship. The current study points in that direction, but longitudinal data would make the inference much stronger.

Reddit is a useful warning against simplistic conclusions

Recent citation volatility also shows why licensing cannot be treated as a permanent visibility guarantee. Reddit has a widely reported commercial relationship with OpenAI, yet ChatGPT citation patterns for Reddit have changed sharply during recent search-system updates. A platform can have privileged access to a source and still alter when or how often it exposes that source in answers.

This suggests that content access and citation preference should be modeled separately. Licensing may increase the availability of content to an AI system without guaranteeing that the content will be surfaced for every relevant query. Retrieval architecture, query fanout, source-quality policies and answer-generation behavior can still reshape the final citation mix.

For unlicensed publishers, the study contains surprisingly good news

The licensing headline can obscure another finding from the research. According to OtterlyAI's summary of the study, trade and niche media outperformed mainstream media in 15 of 16 industries analyzed. Within trade and niche media, licensed publishers accounted for only 17.2% of citations.

That suggests specialized editorial relevance remains a powerful route into AI answers even without a commercial relationship. Unlicensed publishers such as specialist information sites can still become highly visible when their content closely matches the questions AI users ask.

For GEO strategy, that is arguably the most actionable conclusion. Most publishers will never negotiate a licensing agreement with a major AI company. They can, however, build strong topical identity, publish service-oriented content, answer specific questions clearly and become a source that retrieval systems repeatedly encounter when resolving a narrow class of queries.

AI source selection may have more than one layer

The study ultimately challenges the idea that every AI citation can be understood as the output of a single neutral ranking contest among web pages. Generative systems operate across multiple layers: model knowledge, licensed datasets, web crawling, search retrieval, direct feeds, source policies and answer synthesis. Different publishers may enter that pipeline through different routes.

That does not mean citations are for sale. It means the information environment feeding an AI product may not be uniform. Commercial agreements can affect what content a platform is permitted to access and how directly that content is integrated into the product. Whether those access advantages translate into preferential citation is exactly the question that needs independent measurement.

The 48% figure is therefore not the conclusion. It is the beginning of a better experiment.

For the next phase of GEO, we should stop asking only whether a brand or publisher is visible. We should ask whether an AI system appears to prefer certain sources after controlling for the reasons those sources would naturally deserve visibility. If that preference exists — and if commercial content relationships help explain it — AI search optimization will need to account not only for relevance and authority, but also for the architecture of access behind the answer.

0%