Publishers with OpenAI deals get 48% more ChatGPT citations. We do not yet know why.

An analysis of 129.3 million AI citations found that OpenAI-licensed publishers received substantially more ChatGPT citations per cited page than…

Publishers with OpenAI licensing agreements were associated with substantially higher ChatGPT citation rates in this dataset. We do not yet know why.

Author: Ian Ash

Published: August 22, 2026

Category: Research

For the past several years, AI companies and publishers have been negotiating an entirely new kind of commercial relationship. Publishers provide valuable content. AI companies want access to it for training, retrieval, products and answers. Licensing agreements create a mechanism for compensating publishers for some of that access. But those agreements raise another question: does having a commercial relationship with an AI company affect how often that publisher appears in the company's answers?

A newly released analysis from Press Ranger and OtterlyAI suggests there may be a relationship, at least in ChatGPT. The companies analysed 129.3 million AI citations across more than 20 million cited URLs during June 2026 and matched them against 91 confirmed AI-publisher licensing agreements covering 314 publisher domains. Their headline finding: pages from publishers with an OpenAI licensing agreement averaged 10.2 ChatGPT citations per cited page, compared with 6.9 citations per page for publishers without one. That is a 48% difference.[1]

<h2>What the researchers actually studied</h2>

The joint study combines two datasets. OtterlyAI supplied 129.3 million citations involving more than 20 million cited URLs, collected during June 2026 across seven AI-search environments: ChatGPT, Google AI Overviews, Google AI Mode, Perplexity, Microsoft Copilot, Gemini and Claude. Press Ranger maintains an AI-Publisher Licensing Research database and identified 91 confirmed agreements between AI companies and publishers, mapped across 314 publisher domains, with the underlying agreements supported by public sources through July 28, 2026. The researchers then examined whether publishers with licensing agreements were cited differently from publishers without them.[1]

<h2>The OpenAI difference is unusually large</h2>

The association becomes even larger when the researchers isolate publishers that had agreements only with OpenAI. Those publishers reportedly received 112% more ChatGPT citations per cited page than unlicensed publishers. The OpenAI cohort also averaged 10.7 citations per cited page across all seven platforms, compared with 7.3 for unlicensed publishers, a reported difference of 46%. But another result is perhaps more revealing. Publishers with OpenAI agreements received 57.9% of their total observed AI citation volume from ChatGPT alone. The citation mix for unlicensed publishers was more evenly distributed across ChatGPT and Perplexity. In other words, the observed difference was not simply that OpenAI's publisher partners happened to be cited more frequently everywhere. Their citation portfolio was particularly concentrated in ChatGPT.[1]

<h2>The same pattern did not appear consistently elsewhere</h2>

If licensed publishers simply tended to be larger, better or more authoritative publishers, one might expect them to outperform on the platforms of their respective licensing partners more generally. That is not what the researchers report. Publishers with Google licensing agreements were cited at a slightly lower rate than comparable unlicensed publishers in Google AI Overviews. Publishers with Perplexity agreements were essentially at parity with unlicensed publishers on Perplexity. Press Ranger and OtterlyAI therefore describe OpenAI as the only licensor in their analysis showing a clear home-platform advantage.[1]

<h2>The most important word here is association</h2>

There are several possible explanations for the result. One is the obvious one: OpenAI's commercial agreements somehow facilitate greater use or citation of licensed publisher content in ChatGPT. That is possible. But the published analysis does not demonstrate it. Another explanation is selection bias. OpenAI is unlikely to select publisher partners randomly. The publishers attractive enough for OpenAI to license may already differ systematically from unlicensed publishers in ways that also make them more attractive citation sources. They may have larger content archives, higher publishing frequency, greater topical authority, stronger brands, more evergreen content, more product and service journalism, better structured websites, greater domain authority, more comprehensive coverage, or content categories particularly useful to ChatGPT. Any of those factors could contribute to the observed citation difference. Independent coverage of the study has highlighted exactly this limitation: the public release does not provide sufficient controls for factors such as publisher scale or content mix to establish that licensing itself caused the difference.[2]

The responsible conclusion is therefore: publishers with OpenAI licensing agreements were associated with substantially higher ChatGPT citation rates in this dataset. The reason for that association remains an open question.

<h2>The 112% result does not solve the causality problem</h2>

At first glance, the OpenAI-only result looks like stronger evidence. Publishers licensed only to OpenAI received 112% more ChatGPT citations per page than unlicensed publishers, and the same publishers did not show an equivalent premium across the other platforms. That strengthens the case that something ChatGPT-specific may be happening. But it still does not identify what. For example, perhaps OpenAI selected those particular publishers because their content filled specific gaps in ChatGPT's information environment. If so, their higher citation frequency could reflect the same underlying characteristics that made OpenAI want the licensing agreement in the first place. This is a classic causal-inference problem: A and B occur together does not establish that A caused B. The ideal analysis would compare citation behaviour before and after a licensing agreement, against an appropriately matched control group of similar publishers without agreements. The publicly released study does not appear to provide that analysis.[1][2]

<h2>But the study contains another finding that may matter even more for AEO</h2>

The licensing result is the headline. For brands, however, another result may be more actionable. The researchers examined news citations across 16 U.S. industries. In 15 of the 16 industries, niche and trade publications captured the majority of news citations. Those specialist outlets collectively received 213% more AI citations than mainstream media, according to the study. And most of them did not have licensing agreements with AI companies.[1]

That is significant, because one of the assumptions beginning to creep into off-engine AEO is that bigger publisher means better AI source. This study suggests the reality may be considerably more nuanced. Topical authority may matter more than generic prestige. Consider a company trying to establish authority in a specialised B2B category. The obvious PR instinct might be to secure coverage in the biggest possible national publication. But an answer engine trying to construct a detailed category recommendation may prefer a specialist trade publication with deep subject expertise. That creates a different off-engine strategy. Instead of asking only which publishers have the greatest human reach, brands should increasingly ask which publishers AI actually uses as evidence for the particular topic in question. Those lists may be quite different.

<h2>News itself is a relatively small part of the citation universe</h2>

Another useful finding helps put the publisher debate in context. The researchers report that news accounted for only 7.2% of all AI citations in the dataset. So publisher licensing is not synonymous with the entire AI evidence environment. Most citations came from elsewhere. But within licensed-publisher citations, the distribution was highly concentrated. Five media groups, Future plc, Forbes, People Inc., Conde Nast and Hearst, accounted for 69% of citations to licensed publishers. That suggests another form of source concentration inside AI search. A relatively small number of publishing organisations can potentially account for a disproportionate amount of the professionally produced editorial evidence being surfaced by answer engines.[1]

<h2>AI seems to favour particular kinds of publisher content</h2>

The content mix is also revealing. According to the study, 46.9% of citations to licensed publishers came from best-of lists, buying guides and product reviews. That makes intuitive sense. These formats map extremely well onto commercial AI prompts: what is the best laptop for a university student, which CRM should a small business use, what is the best hotel in Rome. The publisher has already done much of the synthesis the answer engine needs. It has identified alternatives, compared them, evaluated attributes and often made a recommendation. For brands, that means the off-engine AEO value of editorial coverage may depend heavily on what kind of content the brand appears in, not simply whether it receives press coverage. A mention in a corporate-news story and inclusion in a respected category buying guide are not necessarily equivalent machine assets.[1]

<h2>This has important implications for off-engine AEO</h2>

The simplistic version of off-engine AEO is: get your brand mentioned on third-party websites. The evidence increasingly suggests a more sophisticated model is needed. The relevant questions become: which sources, not simply high-authority sites but sources the answer engine actually retrieves for the category; which content, since a product review, comparative article or buying guide may have very different machine utility from a general company profile; which engine, since a source influential in ChatGPT may not have the same importance in Gemini or Perplexity; which topic, since a trade publication may dominate evidence for one category while being irrelevant to another; and now, potentially, which relationship, since the publisher may have a licensing or commercial relationship with the answer engine. The causal significance of that final variable remains unresolved. But after this study, it is difficult to argue that it is not worth measuring.

<h2>AEO measurement may need a new piece of metadata</h2>

Most citation analysis currently records the engine, prompt, URL, domain, citation position, brand and date. It may eventually need another attribute: platform-publisher relationship. For example: no known relationship, OpenAI licensed, Google licensed, Perplexity licensed, or multiple platform agreements. Then practitioners can observe whether citation patterns differ systematically. Not because anyone should assume commercial influence, but because if the goal is to understand why the evidence environment behaves the way it does, known structural relationships between platforms and sources are relevant variables.

<h2>This connects to a much larger question about AI search</h2>

Traditional search engines built their source universe largely by crawling the open web. Generative AI is developing in a more complicated information economy. The evidence available to an answer engine can now come through multiple channels: open-web crawling, search indexes, licensed publisher content, structured data feeds, commercial partnerships, APIs, user-generated content, proprietary datasets and potentially other arrangements. Those are not necessarily equivalent forms of access. So when an AI system cites a publisher, the question increasingly becomes: how did that information enter the system's accessible evidence environment? That question is becoming part of AEO.

<h2>But do not turn this into pay to play</h2>

The most irresponsible interpretation of this study would be that publishers can pay or partner their way into ChatGPT citations. The research does not show that. Indeed, the study's own other findings argue against such a simplistic conclusion. Unlicensed niche and trade publications dominated news citations across 15 of 16 industries examined. The most-cited news domains also included unlicensed publications such as NerdWallet, Healthline and Bankrate. So licensing clearly is not a prerequisite for substantial AI visibility. If anything, the results suggest two things may simultaneously be true: OpenAI-licensed publishers exhibit an unusually large ChatGPT citation association, and highly relevant unlicensed specialist publishers remain extremely important sources for AI systems. Those are not contradictory findings.[1]

<h2>AEO Updates Takeaway</h2>

The 48% number is interesting. But the more important development is that commercial relationships are now becoming another observable dimension of the AI evidence ecosystem. That ecosystem is getting increasingly complicated. Practitioners already need to understand source authority, source geography, source concentration, evidence independence, content format, topical relevance, freshness and engine-specific retrieval behaviour. Now they may need to understand platform-source relationships as well. The answer engine does not operate over an abstract, uniform internet. It operates over an information environment shaped by technical access, retrieval systems, publisher economics, content licensing, commercial agreements and, increasingly, optimisation for the answer engines themselves. That is why citation counts alone tell us progressively less. Knowing that a publisher received 10,000 citations is useful. Understanding why that publisher became part of the machine's evidence environment is much more valuable. And this study gives us another reason to start asking that question.

<h3>References</h3>

[1] Press Ranger and OtterlyAI. "AI-Publisher Licensing Study: 129.3 Million Citations Analysed." August 20, 2026. This is the primary source for the reported findings, documenting the 48% difference, 112% OpenAI-only finding, 57.9% ChatGPT concentration, Google and Perplexity comparisons, 7.2% news share, 69% licensed-publisher concentration, 213% niche/trade advantage and 46.9% content-format finding.

[2] Venture Post. "OpenAI Publisher Licensing and ChatGPT Citation Rates." August 21, 2026. Independent coverage that explicitly distinguishes the reported association from causality and notes the absence of controls for confounding variables.

[3] Yahoo Finance / GlobeNewswire. August 20, 2026. Syndicated version of the companies' release corroborating dataset size, platform coverage and major numerical findings.

Continue exploring