Nearly 1 in 11 retrieved pages may already be GEO-optimised

A new detection paper estimates that 8.9 per cent of webpages in released Google Search and Gemini-grounded retrieval results show evidence of Generative…

When machines retrieve information from a web increasingly engineered for machines, how do they distinguish the best evidence from the best-optimised evidence?

Author: Ian Ash

Published: August 19, 2026

Category: Research

Generative Engine Optimisation has moved remarkably quickly from theory to practice. Researchers are now attempting to measure how much of the web has already been optimised for generative engines.

A paper released on 17 August 2026, titled GEO-Flag: Detecting and Measuring GEO-Optimized Web Content, estimates that 8.90 per cent of webpages in a large sample of released Google Search and Gemini-grounded retrieval results showed evidence of GEO optimisation. Among pages modified in 2026, the researchers estimate the prevalence at 16.36 per cent. Put more intuitively, that is roughly one in 11 retrieved pages overall, and roughly one in six recently modified pages.

Those numbers require important qualification. The researchers did not randomly sample the entire web and determine that 8.9 per cent of all webpages have been GEO-optimised. They built a detection system and applied it to 10,095 available webpages retrieved for 1,000 real-user queries in released Google Search and Gemini-grounded retrieval results. The finding therefore concerns the researchers' sampled retrieval environment, not the entire internet.

But even with that qualification, the result raises an important question for AI search: what happens when the evidence an AI system retrieves has itself increasingly been engineered to influence AI systems?

<h2>What the researchers actually did</h2>

The paper was authored by Junjie Chu, Ye Leng, Mingjie Li, Yun Shen, Xinyue Shen and Yang Zhang. Their objective was not primarily to determine which GEO tactics work best. They were trying to solve a different problem: can we reliably detect when a webpage has been deliberately optimised for generative engines?

That is harder than it sounds. A page written clearly, containing statistics, citations and concise explanations might simply be good content. A page edited by AI might also exhibit stylistic characteristics associated with GEO without actually having been optimised for an answer engine. Detecting GEO therefore requires distinguishing among at least three possibilities: ordinary human-authored content, content that has been polished or rewritten using AI, and content deliberately modified to increase its probability of selection or citation by a generative engine.

The researchers created a benchmark called GEOFlagBench containing 3,200 webpages across 400 queries, four domains and eight families of GEO optimisers to evaluate whether those distinctions could be detected reliably. That benchmark is important because the researchers were not simply asking an LLM whether something looks GEO-optimised. They attempted to construct and validate an actual detection methodology.

<h2>Detection is possible, but not trivial</h2>

The researchers tested existing detection approaches and found that the strongest baseline achieved an aggregate F1 score of 0.880. But they also found weaknesses. When they examined performance by optimisation method and authorship conditions, some detection approaches appeared vulnerable to what they describe as authorship-related shortcuts, effectively identifying characteristics associated with how something was written rather than reliably identifying the GEO intervention itself.

They therefore developed a technique called Intervention-Paired Training, or IPT. The important idea is relatively simple. Instead of merely teaching a detector what GEO content tends to look like, IPT attempts to teach it to distinguish specific GEO interventions from ordinary AI-assisted polishing. Using ModernBERT, the researchers report that IPT increased F1 performance from 0.862 to 0.944 and improved worst-group accuracy from 0.725 to 0.883.

That distinction matters enormously. If GEO detection simply becomes another form of AI-content detection, it will not tell us very much. The strategically interesting question is not whether AI was involved in writing a page. It is whether the page was deliberately altered to improve its treatment by generative search systems. The paper represents an early attempt to answer that second question systematically.

<h2>Into the wild: prevalence in real retrieval results</h2>

After developing and testing their detection pipeline, the researchers applied it to released retrieval results associated with 1,000 real-user queries. The resulting dataset contained 10,095 available webpages from Google Search and Gemini-grounded retrieval results. Their estimated GEO prevalence was 8.90 per cent overall, and 16.36 per cent among pages modified during 2026.

Those figures should not be generalised into statements such as "16 per cent of the web is GEO optimised." The study does not establish that. But it does provide evidence that GEO-like optimisation may already be present at meaningful levels among webpages appearing in actual retrieval environments examined by the researchers. That is a considerably more interesting finding.

<h2>From optimisation to competition</h2>

The first stage of GEO was relatively straightforward. Researchers and practitioners asked how a webpage could increase its chances of being cited by a generative engine. The original GEO literature explored interventions such as adding statistics, quotations, citations and other content characteristics intended to improve visibility in generative answers. Commercial AEO platforms have subsequently turned optimisation into workflows involving content structure, entity clarity, citations, authority, source development and increasingly sophisticated measurement.

But once a meaningful portion of publishers begins optimising for machine retrieval, the problem changes. It is no longer optimised page versus unoptimised environment. It becomes optimised page versus other optimised pages competing for machine selection. That is much closer to the competitive dynamic that eventually developed around SEO.

<h2>AI search has an additional problem</h2>

Traditional SEO largely competed for position. GEO potentially competes for something more consequential: evidence selection. A search engine historically presented a collection of competing links and allowed the user to evaluate them. Generative engines increasingly retrieve information from multiple sources and synthesise those sources into a single answer.

The authors of GEO-Flag specifically identify this distinction as a potential risk. They argue that strategically optimised pages can potentially receive visibility disproportionate to their authority or relevance, while weak or false information may appear well supported when generative systems synthesise it into an answer.

That is a fundamentally different optimisation problem. If a weak webpage moves from position nine to position four in traditional search, the user still sees competing sources. If a generative engine selects information from that webpage and incorporates it into a confident synthesised answer, the optimisation can potentially affect what the machine appears to believe.

<h2>The evidence environment is no longer passive</h2>

This creates a recursive problem for AI search. Publishers and brands increasingly know that machines are reading their content. And they can deliberately change the information those machines encounter. The evidence environment is no longer passive. That is precisely why the ability to detect GEO could become important.

Today, AEO platforms largely help brands perform optimisation. Tomorrow, answer engines may increasingly need systems that detect it. That could create a familiar technological cycle: optimisation, then detection, then algorithmic response, then new optimisation techniques, then better detection, then further algorithmic response. SEO experienced versions of this cycle for decades. AI search may be heading toward its own version, but the stakes could be higher because optimisation can potentially affect not merely which page appears first, but which information is incorporated into the generated answer itself.

<h2>What counts as manipulation?</h2>

Not all GEO is undesirable. In fact, much of what gets described as AEO or GEO is simply good information architecture. A company can make its product specifications clearer. A publisher can cite primary sources. An organisation can correct outdated information. An expert can make claims more precise. A webpage can clearly identify authorship. Structured data can help machines understand entities. Original research can provide genuinely useful information. Those are not obviously manipulative behaviours.

The problem begins when optimisation becomes primarily about increasing machine selection without increasing information quality, or worse, when it causes low-authority or poorly supported claims to receive disproportionate visibility. The GEO-Flag researchers therefore go beyond detection. Their system also includes a GEO-gated agent for auditing the source tier and verifiability of citation URLs found in pages detected as GEO-optimised. That is an important conceptual step. The question is not merely whether something was optimised. It becomes whether it was optimised and whether the evidence is actually good.

<h2>A measurement problem for AEO</h2>

The paper creates another interesting challenge for AEO measurement. Suppose a competitor suddenly increases its citation share. The obvious conclusion might be that their content became more authoritative. But perhaps they simply began aggressively optimising their pages for generative retrieval. Or suppose a domain suddenly disappears from AI citations. Perhaps its authority did not decline at all. The answer engine may have changed how it identifies or discounts optimisation.

That means future AEO intelligence may need to distinguish between observed AI visibility and the mechanisms producing that visibility. Those are not the same thing.

<h2>16.36 per cent may be the more interesting number</h2>

The overall 8.90 per cent estimate is the headline. But the 16.36 per cent estimate among pages modified in 2026 may ultimately be more revealing. It suggests, within this particular dataset and detection methodology, that GEO characteristics are substantially more prevalent among recently modified content. That is consistent with what we would expect if publishers and marketers are actively adapting existing content to the generative-search environment.

But this needs to be treated as suggestive rather than causal evidence. The study does not prove that the increase is entirely the result of a sudden industry-wide rush to perform GEO. The estimate depends on the researchers' detector, sample and definition of GEO optimisation. Still, it provides an intriguing empirical signal. The web being retrieved by generative systems may already be changing in response to the existence of generative systems.

<h2>AEO is becoming reflexive</h2>

That may be the biggest implication of the research. AI systems observe the web. Publishers observe what AI systems select. Publishers change the web to improve their chances of selection. AI systems then observe the changed web. Platforms may subsequently change their retrieval systems in response. And publishers optimise again. The system becomes reflexive. AI search changes the web it relies upon. And the changed web subsequently changes AI search. That is the beginning of an optimisation ecosystem.

<h2>AEO Updates Takeaway</h2>

The most important finding in GEO-Flag may not ultimately be whether the true number is exactly 8.90 per cent. The methodology is new. Detection remains imperfect. The dataset is not equivalent to the entire web. And future research will need to test whether these estimates replicate across other engines, query populations, languages, domains and time periods. The bigger development is that researchers now consider GEO sufficiently prevalent to require systematic detection, auditing and measurement. That represents a maturation of the field. The first question was whether content could be optimised for generative engines. The next was whether that optimisation actually increases visibility. Now another question is emerging: can generative engines tell when it is being done? And after that comes the question that may matter most: when machines retrieve information from a web increasingly engineered for machines, how do they distinguish the best evidence from the best-optimised evidence?

<h3>References</h3>

[1] Chu, Junjie; Leng, Ye; Li, Mingjie; Shen, Yun; Shen, Xinyue; Zhang, Yang. "GEO-Flag: Detecting and Measuring GEO-Optimized Web Content." arXiv, submitted 17 August 2026. <a href="https://arxiv.org/abs/2608.16824v1">arxiv.org</a>

Primary sources cited

This article links directly to the primary documentation, paper, filing or original reporting used for its material claims.

  1. arxiv.org

Continue exploring