AI search is generating transactions. We still do not know what visibility is worth

Referral data, consumer experiments and agent simulations show that AI can influence commercial outcomes.

AI search has crossed from hypothetical value into measurable commerce. The unresolved question is how much of that value was caused by the answer, rather than merely observed after it.

Author: Ian Ash

Published: August 29, 2026

Category: Research

What is one additional point of AI visibility worth?

The honest answer is that no credible universal estimate exists.

There is now enough evidence to reject the idea that AI search has no commercial value. Large observational datasets contain AI-referred transactions and revenue. Consumer experiments show that AI recommendations can alter consideration and choice. Agent simulations show that recommendations can change sharply when sources, source order, memory or retrieval paths change.

Those findings matter. They do not solve the valuation problem.

A brand can gain ten points of visibility without gaining recommendation share. It can gain recommendation share without creating a click. It can influence a buyer who later returns directly, searches on Google, visits a retailer or purchases offline. It can also receive highly converting AI referrals because the assistant selected people who were already unusually ready to buy.

The central distinction is between <strong>commercial outcomes that occur after an AI interaction</strong> and <strong>incremental outcomes caused by a change in AI visibility</strong>.

AI search has crossed from hypothetical value into measurable commerce. The unresolved question is how much of that value was caused by the answer, rather than merely observed after it.

<h2>The strongest transaction study finds real revenue at very small scale</h2>

The most comprehensive published transaction evidence comes from Maximilian Kaiser and Christian Schulze. Their peer-reviewed <a href="https://pubsonline.informs.org/doi/10.1287/mksc.2025.0489" target="_blank" rel="noopener noreferrer">Marketing Science study</a> analysed first-party Google Analytics data from 973 e-commerce websites across 24 categories.

Its 12-month robustness dataset contained 10.5 billion sessions, 164.9 million transactions and $20.6 billion in total site revenue. Within that much larger environment, 4.9 million ChatGPT-referred sessions produced 50,251 transactions and approximately $7 million in directly attributed revenue.

The $20.6 billion is not AI revenue. It is the combined revenue observed across the participating sites. The AI-referred component remained very small. Organic large-language-model traffic accounted for less than 0.2% of all traffic and was roughly 200 times smaller than Google organic traffic.

The study’s adjusted six-month models produced a more nuanced result than the claim that AI referrals simply ‘convert better’. Paid-social traffic had a 53% lower conversion likelihood than organic LLM traffic. Organic-search traffic had a 13% higher conversion likelihood than organic LLM traffic. AI referral revenue per session was above paid social and below every other measured traditional channel.

Product complexity moderated the result. Organic LLM financial performance and traffic shares were stronger in more complex categories, consistent with AI being more useful when the buyer needs explanation, comparison or synthesis.

The study is descriptive, not causal. Last-click attribution observes the channel that delivered the visit and purchase. It misses the AI conversation that influenced someone who later arrived through direct traffic, organic search or another path. It can also overstate causal value if people who click from ChatGPT already differ systematically from people arriving through other channels.

The finding is still substantial: <strong>AI referral sessions now produce observable transactions and revenue across a large multi-site dataset. Their measured scale remained tiny during the study period, and their performance sat between paid social and most established channels.</strong>

<h2>Shopify finds stronger lower-funnel performance within a narrower comparison</h2>

Shopify’s <a href="https://www.shopify.com/enterprise/blog/ai-search-insights" target="_blank" rel="noopener noreferrer">Q1 2026 commerce analysis</a> reports a more favourable AI-referral pattern.

Among sessions beginning on product-detail pages, AI-referred visits converted 49% more often than organic-search visits. The advantage appeared in 23 of 25 merchant categories and averaged 56% within those categories. Orders attributed to AI-powered search also had 14% higher average order value.

The denominator is important. The 49% result compares <strong>product-detail-page entry sessions</strong>, not all AI and organic traffic. Shopify reported that 55% of AI-referred sessions began on a product page, compared with about 20% of organic-search sessions. The groups therefore arrive with different journey histories and likely different intent.

Shopify also reported more than eightfold year-over-year growth in AI referral sessions and nearly thirteenfold growth in AI-referred orders. Organic search still delivered more sessions than all tracked AI platforms combined. Shopify further notes that referrals from Google AI Overviews can be classified as organic search, making clean channel separation difficult.

This is aggregated commercial platform data rather than a peer-reviewed study, and Shopify’s public article does not disclose the underlying merchant, session or order counts. It supports the conclusion that some AI-referred Shopify sessions are commercially valuable. It does not show that a visibility gain caused the conversion difference.

<h2>Adobe shows how channel performance can change with time and category</h2>

Adobe Digital Insights analysed more than one trillion visits to US retail websites and paired that behavioural dataset with a survey of 5,000 US consumers. Its <a href="https://blog.adobe.com/en/publish/2025/03/17/adobe-analytics-traffic-to-us-retail-websites-from-generative-ai-sources-jumps-1200-percent" target="_blank" rel="noopener noreferrer">March 2025 report</a> found that generative-AI referral traffic in February 2025 was 1,200% above July 2024.

The growth was dramatic, but Adobe described the channel as modest compared with paid search and email. AI-referred visitors showed 8% higher engagement, viewed 12% more pages and had a 23% lower bounce rate than non-AI traffic. They were also 9% less likely to convert, an improvement from a 43% deficit in July 2024.

Electronics and jewellery had the highest AI-referral conversion rates in Adobe’s category comparison, while apparel, home goods and grocery were lower. That pattern again suggests that the value of AI referrals may vary with information complexity and decision structure.

Adobe, Shopify and Kaiser and Schulze are not contradictory datasets. They cover different periods, sites, channel definitions and comparison bases. Together they make a stronger point: <strong>there is no universal ‘AI traffic conversion rate’.</strong> Category, entry page, time, attribution method and user selection all matter.

<h2>Consideration can move before a transaction appears</h2>

Referral analytics begin when someone clicks. Much of the potential value may occur earlier, when AI changes which brands or products enter the consideration set.

A 2024 <a href="https://doi.org/10.1016/j.jretconser.2024.103743" target="_blank" rel="noopener noreferrer">Journal of Retailing and Consumer Services study</a> recruited 471 US consumers through Amazon Mechanical Turk and retained 443 valid responses. In a within-subject design, participants completed one product-recommendation mission with Amazon and another with BingChat as a ChatGPT proxy.

Both systems were associated with consideration-set intention through trust in the recommender and the recommended product. BingChat produced higher reported consideration-set intention and a stronger perceived-performance-to-trust path. The trust gap narrowed for products with low and medium brand awareness, but not for high-awareness products.

The outcome was intention to adopt a product into a consideration set. It was not an actual transaction. The two interfaces and product missions also differed. The study nevertheless establishes a commercially relevant stage that referral reporting misses: AI can affect which options people say they would keep under consideration.

This is why the AEO Updates framework separates <a href="/what-is-aeo">Mention, Consideration, Recommendation and downstream Choice or behaviour</a>. A mention is not a shortlist. A shortlist is not a preferred recommendation. A preferred recommendation is not a purchase.

<h2>Recommendation position can influence scenario choice</h2>

Placement inside an AI response may matter independently of whether the brand appears at all.

A 2026 <a href="https://doi.org/10.1016/j.elerap.2026.101605" target="_blank" rel="noopener noreferrer">Electronic Commerce Research and Applications paper</a> ran three scenario-based experiments involving product and travel recommendations. Participants displayed a robust first-item preference in ChatGPT lists even when the researchers randomised item order, removed numerical ranking cues or placed an objectively flawed option first.

Decision-time evidence was consistent with lower-effort, position-based processing. Participants primed towards prevention were less likely than promotion-focused participants to rely on the first item and spent longer deliberating.

These were scenario choices, not real purchases. The study cannot assign a revenue value to first position. It does show why ‘the brand appeared’ is economically incomplete. Recommendation order and presentation can alter the choice made inside an experiment even when no explicit ranking rationale is visible.

<h2>AI assistance may reduce the protection offered by reputation</h2>

An August 2026 <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7250218" target="_blank" rel="noopener noreferrer">SSRN working paper by Louis Krol and Simone Santamaria</a> tested a different mechanism.

In a randomised online experiment, 350 participants made 1,750 incentivised choices across Bluetooth speakers, computer mice, disposable razors, electric toothbrushes and swimming goggles. Some participants had access to an embedded generative-AI assistant. A lottery awarded selected products or their cash equivalent.

In the control condition, participants chose the more reputable but attribute-weaker brand in 37.3% of choices. With AI access, that share fell to 23.8%, a 13.5 percentage-point difference. AI participants also reported greater reliance on functional or technical specifications, 48.5% versus 40.0%, and less reliance on brand name, 14.6% versus 20.3%.

The paper is recent and not peer reviewed. Lottery-backed choices are stronger evidence than a hypothetical intention scale, but they are not purchases made with participants’ own money. The five categories were also relatively objective and specification-heavy.

The result should not be generalised into ‘AI kills brands’. It supports a narrower hypothesis: when AI lowers the cost of comparing functional attributes, reputation may provide less protection to an established brand whose product is weaker on the displayed criteria.

That mechanism connects directly to <a href="/articles/what-makes-brand-show-up-ai-evidence-environment">Machine Positioning</a>. A brand needs more than recognition. It needs a credible and retrievable connection to the attributes that matter in the decision.

<h2>Agent recommendations are themselves unstable</h2>

The recommendation probability on the AI side of the value chain is not fixed.

The Wharton Generative AI Labs <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7355899" target="_blank" rel="noopener noreferrer">Research Report 6</a> used a simulated fitness-watch store to test more than 26,000 product selections across six frontier models. The researchers changed the prior sources shown to the models, the order of those sources, user-memory statements and the way retrieval tools delivered information.

Recommendations changed materially across those conditions. Multiple sources did not reliably average out. The same sources in a different order altered selections for some models. Remembered preferences and retrieval paths also changed choices, sometimes even when one product was objectively stronger on price, rating and review count.

This is a model-behaviour experiment in a fixed simulated store, not evidence that human consumers purchased the selected products. It does reveal a hidden variable in AEO valuation. The likelihood of recommendation can change with model, source mix, source order, memory and retrieval path before any human response is observed.

That instability is one reason a single prompt snapshot cannot support a revenue claim.

<h2>One agency case study reaches beyond referrals, but not causality</h2>

Commercial case studies are beginning to combine prompt tracking, referral revenue and customer surveys.

Visibility Labs reports that a six-month programme for Private Label MFG increased recommendation visibility across 100 selected prompts from approximately 1% to more than 20%. The agency also reports citation rate rising from 7% to 18.4%, AI referral revenue increasing 344%, ChatGPT accounting for 94.6% of AI referral sessions and post-purchase survey responses naming an AI assistant as the discovery channel rising from 0.5% to 5%.

The case is useful because it attempts to observe more than one stage. It is also low-confidence evidence. The agency designed the programme and published the result. There was no control group. Website changes, content production, public mentions and Reddit activity changed at the same time. GA4 captured direct referrals, while the survey relied on customer recall. The agency page and a later press release also give inconsistent campaign years.

The result shows how one company assembled an attribution stack. It does not show that visibility growth caused the reported revenue growth, or that the same programme would produce similar results elsewhere.

<h2>The missing research question is incrementality</h2>

The current evidence answers several important questions.

<div class="my-8 overflow-x-auto"><table class="w-full min-w-[840px] border-collapse text-sm"><thead><tr class="border-b-2 border-[#20160E]"><th class="px-3 py-3 text-left font-semibold">Question</th><th class="px-3 py-3 text-left font-semibold">Best available evidence</th><th class="px-3 py-3 text-left font-semibold">What remains unresolved</th></tr></thead><tbody><tr class="border-b border-[#D7C8B5]"><td class="px-3 py-3 font-semibold">Do AI referrals produce transactions?</td><td class="px-3 py-3">50,251 observed ChatGPT-referred transactions</td><td class="px-3 py-3">How many were incremental rather than journeys that would have converted elsewhere</td></tr><tr class="border-b border-[#D7C8B5]"><td class="px-3 py-3 font-semibold">Can AI affect consideration?</td><td class="px-3 py-3">Higher consideration-set intention in the BingChat condition</td><td class="px-3 py-3">Real-market purchasing and long-run brand effects</td></tr><tr class="border-b border-[#D7C8B5]"><td class="px-3 py-3 font-semibold">Can position affect choice?</td><td class="px-3 py-3">Robust first-item preference across three scenarios</td><td class="px-3 py-3">Economic value, category durability and live-interface effects</td></tr><tr class="border-b border-[#D7C8B5]"><td class="px-3 py-3 font-semibold">Can AI change brand reliance?</td><td class="px-3 py-3">13.5-point drop for the reputable but attribute-weaker option</td><td class="px-3 py-3">Replication across subjective, emotional and high-stakes categories</td></tr><tr class="border-b border-[#D7C8B5]"><td class="px-3 py-3 font-semibold">Can the information environment change an agent’s selection?</td><td class="px-3 py-3">More than 26,000 Wharton simulation runs</td><td class="px-3 py-3">Human response and stable real-world retrieval mechanisms</td></tr><tr><td class="px-3 py-3 font-semibold">Can visibility, traffic and survey attribution move together?</td><td class="px-3 py-3">One Visibility Labs case study</td><td class="px-3 py-3">Causal lift and external validity</td></tr></tbody></table></div>

The most important unanswered question is:

<blockquote>How does an incremental change in AI visibility, representation or recommendation affect incremental business outcomes for a specific brand?</blockquote>

Answering it requires experimentation or a defensible quasi-experimental design. Brands could test markets, time periods, products or prompt clusters, provided the intervention is defined in advance and other changes are controlled as far as possible. The field still lacks enough published evidence to provide a universal response curve.

<h2>An AEO value chain is more honest than an ROI equation</h2>

The useful part of an AEO Valuation Model is the sequence of measurable transitions. Presenting them as one equation would imply coefficients that do not yet exist.

<blockquote><strong>Relevant AI decision occasions → Brand exposure → Accurate representation → Consideration → Recommendation prominence → Verification or handoff → Choice → Conversion → Customer economic value</strong></blockquote>

Each arrow asks a different question and requires a different denominator.

<div class="my-8 overflow-x-auto"><table class="w-full min-w-[900px] border-collapse text-sm"><thead><tr class="border-b-2 border-[#20160E]"><th class="px-3 py-3 text-left font-semibold">Stage</th><th class="px-3 py-3 text-left font-semibold">Example metric</th><th class="px-3 py-3 text-left font-semibold">Required denominator</th><th class="px-3 py-3 text-left font-semibold">Evidence source</th></tr></thead><tbody><tr class="border-b border-[#D7C8B5]"><td class="px-3 py-3 font-semibold">Relevant decision occasions</td><td class="px-3 py-3">Estimated prompt or task demand</td><td class="px-3 py-3">Declared audience, category, geography and time window</td><td class="px-3 py-3">Search data, surveys, behavioural panels or platform data</td></tr><tr class="border-b border-[#D7C8B5]"><td class="px-3 py-3 font-semibold">Brand exposure</td><td class="px-3 py-3">Visibility or Share of Visibility</td><td class="px-3 py-3">Stable prompt universe and named AI system</td><td class="px-3 py-3">Repeated response observation</td></tr><tr class="border-b border-[#D7C8B5]"><td class="px-3 py-3 font-semibold">Accurate representation</td><td class="px-3 py-3">Claim and association accuracy</td><td class="px-3 py-3">Responses in which the brand appears</td><td class="px-3 py-3">Coded answer content against verified brand facts</td></tr><tr class="border-b border-[#D7C8B5]"><td class="px-3 py-3 font-semibold">Consideration</td><td class="px-3 py-3">Consideration-set rate</td><td class="px-3 py-3">Decision prompts where the brand is eligible</td><td class="px-3 py-3">Response coding or consumer research</td></tr><tr class="border-b border-[#D7C8B5]"><td class="px-3 py-3 font-semibold">Recommendation prominence</td><td class="px-3 py-3">Preferred-recommendation or first-position rate</td><td class="px-3 py-3">Recommendation prompts</td><td class="px-3 py-3">Response coding with position</td></tr><tr class="border-b border-[#D7C8B5]"><td class="px-3 py-3 font-semibold">Verification or handoff</td><td class="px-3 py-3">Source click, branded search, direct visit or assisted path</td><td class="px-3 py-3">Exposed or recommended journeys</td><td class="px-3 py-3">Platform, browser, referral and survey data</td></tr><tr class="border-b border-[#D7C8B5]"><td class="px-3 py-3 font-semibold">Choice</td><td class="px-3 py-3">Selected brand or product</td><td class="px-3 py-3">Completed decision occasions</td><td class="px-3 py-3">Experiment, panel, survey or transaction system</td></tr><tr class="border-b border-[#D7C8B5]"><td class="px-3 py-3 font-semibold">Conversion</td><td class="px-3 py-3">Purchase, lead or qualified action</td><td class="px-3 py-3">Measured visits or eligible users</td><td class="px-3 py-3">Analytics, CRM and commerce data</td></tr><tr><td class="px-3 py-3 font-semibold">Customer economic value</td><td class="px-3 py-3">Margin, retention or lifetime value</td><td class="px-3 py-3">Converted customers</td><td class="px-3 py-3">Finance, CRM and cohort analysis</td></tr></tbody></table></div>

The strongest current studies illuminate different rows. None measures the entire chain for the same population, category and period. Multiplying averages from Kaiser, Shopify, Chang, Hong, Krol and Wharton would create false precision.

The <a href="/articles/ai-visibility-no-common-denominator">denominator rule</a> still applies: the numerator cannot be interpreted until the eligible population is visible.

<h2>What marketers should measure now</h2>

The immediate task is not to assign a universal dollar value to one visibility point. It is to improve the organisation’s measurement architecture.

First, maintain two distinct scorecards. The diagnostic scorecard should track Access, Visibility, Brand Representation, Consideration and Recommendation across a stable prompt universe and named AI systems. The commercial scorecard should track direct referrals, assisted discovery, leads, purchases, revenue and customer value. One should not be presented as a substitute for the other.

Second, segment by decision purpose. AI appears more commercially relevant when the task requires research, comparison or synthesis. The <a href="/articles/people-use-ai-search-narrow-decisions">AI decision-purpose analysis</a> provides a practical framework for Learn, Discover, Compare, Narrow, Recommend, Validate and Act contexts.

Third, preserve the handoff. Use referral parameters where available, separate AI referrers in analytics, capture branded-search and direct-traffic movement cautiously, and ask post-purchase discovery questions without forcing AI as an answer. Self-report and web analytics have different biases; convergence is more informative than either alone.

Fourth, record the intervention. If product pages, third-party evidence, content, technical access and prompt coverage all change simultaneously, later attribution will be difficult. A phased design creates a better chance of estimating incremental effects.

Finally, report uncertainty. AI-referred transactions are measured. Unclicked influence is partly hidden. Recommendation volatility is observed in experiments. The causal value of one visibility point remains unknown. An executive report should say all four things.

<h2>AEO Updates Takeaway</h2>

AI search now generates measurable transactions. A peer-reviewed study observed more than 50,000 ChatGPT-referred purchases and almost $7 million in attributed revenue. Shopify and Adobe provide additional, category-sensitive commercial signals. Experiments show that AI can affect consideration, recommendation processing and brand choice, while agent simulations show that the recommendations themselves can be unstable.

That is enough to justify disciplined AEO measurement. It is not enough to justify a universal ROI claim.

The commercial question should no longer be ‘Does AI search create value?’ It sometimes does, and that value can now be observed. The better question is: <strong>which part of the decision journey changed, for which audience and category, and how much of the downstream outcome was incremental?</strong>

Until that chain is measured, a visibility score remains a diagnostic signal rather than a financial result.

<h3>References</h3>

[1] <a href="https://pubsonline.informs.org/doi/10.1287/mksc.2025.0489" target="_blank" rel="noopener noreferrer">Maximilian Kaiser and Christian Schulze, ‘Frontiers: ChatGPT Referrals to E-Commerce Websites: How Do LLMs Compare Against Traditional Channels?’</a>, <em>Marketing Science</em>, 21 April 2026.

[2] <a href="https://www.shopify.com/enterprise/blog/ai-search-insights" target="_blank" rel="noopener noreferrer">Kyle Risley, ‘AI-referred shoppers convert better and spend more: What Shopify’s early data shows’</a>, Shopify, 11 May 2026.

[3] <a href="https://blog.adobe.com/en/publish/2025/03/17/adobe-analytics-traffic-to-us-retail-websites-from-generative-ai-sources-jumps-1200-percent" target="_blank" rel="noopener noreferrer">Vivek Pandya, ‘Traffic to U.S. retail websites from Generative AI sources jumps 1,200 percent’</a>, Adobe Digital Insights, 17 March 2025.

[4] <a href="https://doi.org/10.1016/j.jretconser.2024.103743" target="_blank" rel="noopener noreferrer">Woondeog Chang and Jungkun Park, ‘A comparative study on the effect of ChatGPT recommendation and AI recommender systems on the formation of a consideration set’</a>, <em>Journal of Retailing and Consumer Services</em>, May 2024.

[5] <a href="https://doi.org/10.1016/j.elerap.2026.101605" target="_blank" rel="noopener noreferrer">Suji Hong and Jennifer Hyun Kim, ‘Heuristic or systematic? Understanding consumer information processing of ChatGPT recommendations’</a>, <em>Electronic Commerce Research and Applications</em>, July to August 2026.

[6] <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7250218" target="_blank" rel="noopener noreferrer">Louis Krol and Simone Santamaria, ‘AI-Empowered Customers and the Erosion of Brand Power’</a>, SSRN working paper, 13 August 2026.

[7] <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7355899" target="_blank" rel="noopener noreferrer">Anushka Kumar et al., ‘Research Report 6: Agentic Shopping is Complicated and Contingent’</a>, Wharton Generative AI Labs, revised 27 August 2026.

[8] <a href="https://visibilitylabs.ai/ai-seo-case-study-private-label-mfg/" target="_blank" rel="noopener noreferrer">Jeff Oxford, ‘How We Grew Private Label MFG’s AI Search Visibility from 1% to 20% in 6 Months’</a>, Visibility Labs, 10 April 2026.

Primary sources cited

This article links directly to the primary documentation, paper, filing or original reporting used for its material claims.

  1. Marketing Science study
  2. Q1 2026 commerce analysis
  3. March 2025 report
  4. Journal of Retailing and Consumer Services study
  5. Electronic Commerce Research and Applications paper
  6. SSRN working paper by Louis Krol and Simone Santamaria
  7. Research Report 6
  8. Jeff Oxford, ‘How We Grew Private Label MFG’s AI Search Visibility from 1% to 20% in 6 Months’

Continue exploring