Two 2026 papers disagree on how brittle AI brand visibility is. They converge on a more important point: one response is not a defensible proxy for a…
Ask an AI assistant for the “best CRM”, then ask it to “recommend a CRM” or identify the best option for a 40-person SaaS company. The prompts occupy the same broad decision territory, but they need not produce the same brands.
That is not merely an interesting feature of generative AI. It is a measurement problem.
If a brand appears for one formulation and disappears for another, what does it mean to report that the brand has 50% AI visibility? And if a platform tracks a fixed list of prompts, how much of the result reflects the brand’s position in the market rather than the questions selected by the measurement system?
Two new working papers offer different estimates of the effect. One finds substantial changes in recommendation sets after seemingly modest rewording. The other finds that visibility changes are structured, with the largest differences concentrated among prompts that have moved further apart in meaning. Their methods and outcomes differ, so their figures should not be combined. But both challenge the idea that one canonical prompt can represent a buyer need.[1] [2]
The practical conclusion is not that prompt tracking is useless. It is that <strong>the prompt is an observation, not the market</strong>.
<h2>The first uncertainty exists before the wording changes</h2>
The strongest controlled evidence comes from a May 2026 arXiv preprint by Will Jack, Noah Lehman, Keller Maloney and Sarah Xu. The researchers ran approximately 6,000 paraphrase tests and approximately 6,000 same-prompt rerun controls across OpenAI and Anthropic production models in commercial recommendation tasks.[1]
They measured overlap between recommendation sets using the Jaccard index. A score of 1 means the two lists contain the same brands; a score of 0 means they share none.
Identical prompts were not perfectly reproducible. Same-prompt reruns produced recommendation-set similarity between 0.50 and 0.61, depending on the model and condition. One prompt asked once was therefore not a stable answer even before its wording changed.[1]
Cosmetic rewordings produced a lower average similarity of 0.288, with a clustered 95% confidence interval of 0.215 to 0.361. Rewordings that added constraints, including region, language or specificity, produced 0.135 similarity, with a 95% confidence interval of 0.098 to 0.175.[1]
These are substantial differences in the tested recommendation tasks. But they should remain bounded to an unreviewed preprint, two provider families and the authors’ prompt-construction choices. They do not establish a universal turnover rate for every AI system or category.
The measurement implication is nevertheless important. AI visibility contains at least two sources of variability.
<div class="overflow-x-auto"><table><thead><tr><th>Source of variation</th><th>What changes</th><th>What it means</th></tr></thead><tbody><tr><td>Run-to-run variability</td><td>The answer changes while the prompt stays fixed</td><td>One response is not a stable estimate</td></tr><tr><td>Prompt-expression variability</td><td>The wording changes around an intended need</td><td>A canonical prompt may overfit one expression</td></tr></tbody></table></div>
Both must be measured before a visibility score can be interpreted confidently.
<h2>A second paper finds a more structured effect</h2>
A commercially affiliated SSRN working paper by Jan Ehrlinspiel, Malte Landwehr and Tomek Rudzki reaches a more qualified conclusion. All three authors are affiliated with Peec AI, and the research was conducted through the company’s internal research programme.[2]
The paper uses two observational designs. The first includes 288 human-written prompts across headphones and design agencies, four answer engines and 15 days, producing 17,280 prompt-engine-date answer runs. The second includes 54 base prompts and 1,412 variants across 18 subverticals, five sectors and two answer engines over seven days, producing 20,524 runs.[2]
Its outcome is not recommendation-set overlap. It is the probability that a reference brand appears in an answer.
The largest visibility differences were associated with the lowest semantic-similarity regions. In the human-prompt study, the 0.35 to 0.39 similarity bin was associated with approximately 2.02 to 2.43 percentage points lower visibility than near-identical prompt pairs, against an average visibility rate of about 4.9%. Prompt format also shifted the baseline, while the clearest controlled-study sensitivity appeared in unbranded middle-funnel commercial prompts.[2]
These are conditional associations, not randomised causal effects. The authors describe the findings as exploratory and preliminary. The study does not support a universal semantic-similarity threshold that all measurement systems can apply.
It does support a more useful principle: human wording variation is structured rather than infinite. A measurement programme can sample semantic neighbourhoods, formats, journey stages and engines instead of trying to enumerate every possible sentence.
<h2>The papers disagree less than their headline numbers suggest</h2>
The Jack paper asks how much two recommendation lists overlap. The Peec paper asks how brand-presence probability changes as prompts move away from a reference prompt. The studies also differ in categories, engines, observation windows, prompt construction and statistical design.
Their estimates are therefore not competing measurements of one common effect size.
The responsible cross-study synthesis is narrower: identical reruns can vary; wording and semantic distance can change observed brand presence; prompt format and purpose can shift the baseline; audience, geography and commercial constraints can legitimately change which brands qualify; and one prompt cannot stand in for the full decision space.
That distinction matters because not every difference is measurement noise. “Best CRM for a small business” and “Top CRM for a small company” may be reasonable near neighbours. “Best CRM for a startup” and “Best CRM for an enterprise” describe materially different audiences. Averaging both kinds of variation together destroys information.
<h2>Winning one prompt is not the same as covering a decision space</h2>
A larger Semrush and Kevin Indig study offers a topic-level view. It analysed 1,094 US categories in ChatGPT from January through June 2026. Each category contained five representative questions covering definition, comparison, alternatives, use case and buying.[3]
Under the study’s definition, a category owner had to lead share of mentions, appear in at least four of five prompts and hold a lead of at least five percentage points over the runner-up. Only 15.2% of categories met that threshold. Another 31.2% had an emerging leader, while 53.7% were unsettled.[3]
This is not a paraphrase experiment. The five prompts intentionally represent different questions within a topic. That is precisely why it complements the two wording studies: it tests consistency across a broader decision space rather than robustness around one narrowly defined need.
<div class="overflow-x-auto"><table><thead><tr><th>Measurement layer</th><th>Question answered</th><th>What should remain controlled</th></tr></thead><tbody><tr><td>Prompt observation</td><td>Did the brand appear in this answer?</td><td>Exact prompt, model, mode, date and run</td></tr><tr><td>Near-neighbour family</td><td>Does the brand survive reasonable rewording of the same need?</td><td>Purpose, audience and material constraints</td></tr><tr><td>Decision-space coverage</td><td>Where does the brand appear across related buyer questions?</td><td>Purpose, journey stage, audience, geography and constraint strata</td></tr><tr><td>System condition</td><td>Does the result change by engine or configuration?</td><td>Platform, model, reasoning mode and retrieval availability</td></tr></tbody></table></div>
These layers should not be collapsed into one score before each remains separately visible.
<h2>Purpose can change what the system is being asked to do</h2>
Another Semrush and Kevin Indig study examined 3,981 domain appearances generated from 115 prompts across 14 countries and four AI search surfaces.[4]
The study distinguished a citation, where a domain appears as a source link, from a mention, where the brand name appears in the response. It reported that 61.7% of appearances were citations without a brand mention, 13.2% were both cited and mentioned, and 25.1% were mentions without a citation.[4]
Prompt purpose was associated with sharply different outcomes. Informational prompts had an 89.3% citation rate but an 18% brand-mention rate. Comparative prompts produced a 43.3% mention rate, approximately 2.4 times the informational rate.[4]
The sample is small and prompt length, purpose, engine, country and construction can vary together. The results do not prove that changing one phrase independently caused the difference. They do show why an informational question and a recommendation question should not be treated as interchangeable observations.
<h2>The system configuration can change the evidence path</h2>
Prompt choice is not the only source of variability. A separate Semrush and Kevin Indig comparison ran 100 prompts representing 20 buyer journeys through GPT-5.2 in Instant and Thinking modes.[5]
Only 25.6% of cited domains overlapped between the two modes. Citation rate increased from 50% to 68%, average sources per cited response from 2.6 to 4.5, and reported internal sub-queries by approximately 4.6 times.[5]
This is a 200-response comparison involving one model family and two configurations, not a universal estimate. It demonstrates why the platform and mode must remain visible in the measurement record.
Google also states that AI Overviews and AI Mode may use query fan-out, issuing multiple related searches across subtopics and data sources when developing a response.[6] That official documentation supports the mechanism in Google Search, but it does not reveal every hidden query, quantify the effect of rewording or establish that every AI assistant works identically.
Fan-out therefore helps explain how a prompt can lead to a different evidence path. It is not a substitute for measuring the wider set of questions a consumer might ask.
<h2>A better unit: the prompt family</h2>
The immediate alternative to one-prompt measurement is not thousands of randomly generated questions. It is a structured prompt family.
<div class="my-8 border-l-4 border-[#5B45D6] bg-[#F0EEFF] px-6 py-5"><strong>AEO Updates working definition:</strong> A prompt family is a controlled set of semantically close prompts intended to represent one defined information need while holding purpose, audience, geography and material constraints constant.</div>
The variants inside that family act as near neighbours. Their job is to estimate local robustness: does the brand persist when the same need is expressed in reasonable alternative ways?
The family can then sit inside a broader decision-space map that keeps materially different purposes and conditions separate. A CRM measurement programme might distinguish learning, exploration, comparison, recommendation and selection questions, then stratify further by business size, industry, geography or price constraint.
This structure avoids two opposite errors. One canonical sentence is too narrow. A large unstructured prompt list is too incoherent.
<h2>What should be averaged, and what should stay separate?</h2>
Near-neighbour variants that genuinely preserve the same construct may be treated as repeated observations around that construct. Their role is to estimate robustness.
Purpose, audience, geography, material constraints, journey stage and AI engine should remain explicit dimensions. Their role is to reveal conditional relevance and platform effects.
<div class="overflow-x-auto"><table><thead><tr><th>Dimension</th><th>Default treatment</th><th>Reason</th></tr></thead><tbody><tr><td>Same-prompt reruns</td><td>Summarise as a distribution</td><td>Measures stochastic variability</td></tr><tr><td>True near neighbours</td><td>Aggregate within a declared family</td><td>Estimates wording robustness</td></tr><tr><td>Purpose or journey stage</td><td>Keep separate</td><td>Represents a different decision task</td></tr><tr><td>Audience, geography or material constraint</td><td>Keep separate</td><td>Changes eligibility and fit</td></tr><tr><td>Engine, model or mode</td><td>Keep separate before any roll-up</td><td>Captures system effects</td></tr><tr><td>Citation, mention and recommendation</td><td>Keep separate</td><td>These are different outcomes</td></tr></tbody></table></div>
Any overall index should be a transparent secondary view, not the only view.
<h2>The denominator must describe the sampled opportunity space</h2>
For one defined prompt family, a simple unweighted visibility measure could be expressed as:
<div class="my-8 rounded-xl border border-[#CAD7D3] bg-[#FFF8EB] px-6 py-5 text-center"><strong>Brand appearances ÷ eligible response opportunities</strong></div>
The denominator must disclose the number of prompt variants, repeated runs, engines, modes and observation dates included. It must also state which outcomes qualify as an appearance.
This differs fundamentally from saying that a brand appeared in ten of 20 prompts without explaining how the prompts represent the underlying need.
Weighting may eventually be useful, but it requires evidence. Purpose, audience value, prompt prevalence, platform usage and commercial importance are not interchangeable. Until a defensible weighting basis exists, platforms should provide a transparent unweighted view and keep any planning weights labelled as assumptions.
<h2>Brand survivorship is a useful diagnostic, not yet a standard metric</h2>
The Prompt Problem suggests a richer question than whether a brand appeared once: <strong>Across the relevant prompt families, decision stages and system conditions, how consistently does the brand remain in the consideration set?</strong>
AEO Updates uses <strong>brand survivorship</strong> as a working diagnostic for that persistence. It is not a validated industry metric and should not be converted into a proprietary score without published construction rules and stability testing.
A brand can be robust within one narrow family but cover little of the broader decision space. Another can have wide coverage but weak repeatability within individual families. Reporting robustness and coverage separately preserves that difference.
<h2>A minimum defensible measurement protocol</h2>
A practical programme can begin with seven steps.
<div class="my-8 rounded-xl border border-[#CAD7D3] bg-[#FFF8EB] px-6 py-5"><ol class="list-decimal space-y-2 pl-5"><li>Define the underlying consumer decision or information need.</li><li>Separate material strata such as purpose, audience, geography, journey stage and constraints.</li><li>Build controlled prompt families within each stratum.</li><li>Run every prompt repeatedly on named engines and configurations.</li><li>Record citations, mentions, consideration and recommendation as separate outcomes.</li><li>Report family-level distributions, decision-space coverage and uncertainty before any roll-up.</li><li>Version the prompt register, models, dates and weighting assumptions so changes remain auditable.</li></ol></div>
This does not solve every problem. The field still lacks validated answers for how many near neighbours are enough, how densely to sample a decision space, how many reruns each engine requires and how to estimate real prompt prevalence.
Those are research questions, not reasons to fall back to one sentence.
<h2>AEO Updates Takeaway</h2>
The studies do not produce one universal estimate of prompt sensitivity. They use different outcomes and cannot be pooled. They do converge on a more consequential conclusion: <strong>one prompt is not a defensible unit of truth for AI visibility</strong>.
The next generation of measurement should treat an individual prompt as one observation inside a declared sampling frame. Near neighbours can test robustness. Prompt families can represent defined needs. A broader decision-space map can preserve purpose, audience and constraint differences. Repeated runs and separate engine reporting can show how much uncertainty comes from the system itself.
The practical question is no longer simply, “Did the brand appear for this prompt?”
It is: <strong>Across the realistic ways people express and navigate this decision, how likely is the brand to enter, and remain in, the AI consideration set?</strong>
That is the Prompt Problem.
<h3>References</h3>
[1] <a href="https://arxiv.org/abs/2605.27440" target="_blank" rel="noopener noreferrer">Will Jack, Noah Lehman, Keller Maloney and Sarah Xu, “Paraphrase Brittleness in Production Retrieval-Augmented Commercial Recommendation”</a>, arXiv preprint, 22 May 2026.
[2] <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6914539" target="_blank" rel="noopener noreferrer">Jan Ehrlinspiel, Malte Landwehr and Tomek Rudzki, “Prompt Tracking Works, But Not as a One-Prompt Measurement System”</a>, SSRN working paper, 2 July 2026.
[3] <a href="https://www.semrush.com/blog/chatgpt-topic-authority-study/" target="_blank" rel="noopener noreferrer">Semrush and Kevin Indig, “AI visibility is a topic-level game”</a>, 20 July 2026.
[4] <a href="https://www.semrush.com/blog/the-ghost-citations-study/" target="_blank" rel="noopener noreferrer">Semrush and Kevin Indig, “Why 62% of AI citations don’t lead to brand mentions”</a>, 9 June 2026.
[5] <a href="https://www.semrush.com/blog/chatgpt-reasoning-ai-visibility/" target="_blank" rel="noopener noreferrer">Semrush and Kevin Indig, “Only 25% of cited sources overlap between ChatGPT’s different reasoning modes”</a>, 30 June 2026.
[6] <a href="https://developers.google.com/search/docs/appearance/ai-features" target="_blank" rel="noopener noreferrer">Google Search Central, “AI features and your website”</a>, accessed 30 August 2026.