Every AI Search measurement reflects a prompt set, measurement frame, observation method, evidence record and calculation.
Google Search Console has exposed an awkward truth about AI-search measurement. The platforms do not necessarily give publishers and brands enough data to reconstruct what happened. Google can isolate generative-AI visibility in one report and show query text in another, but it still cannot reliably connect the question, the conversation and the cited page in a single view.[1] [2]
That gap extends beyond Google. Every AEO dashboard begins with an observation. A measurement system has to define a prompt set, collect responses from an AI system, retain raw observations and turn them into explicit measures such as Visibility, Share of Visibility, Citation Rate, Response Claim Frequency or Association Frequency.
The finished dashboard can look precise. The measurement beneath it is shaped by a chain of practical choices. Those choices are not defects. They are the design of the measurement system.
Understanding that design does not require every prompt to come from a perfect dataset or every observation to reproduce an ordinary user. It requires a clearer connection between the purpose of the programme, the way it was constructed and the conclusions the results can support.
<h2>Every dashboard starts with a designed observation</h2>
Suppose a company wants to measure its visibility for the question, “What are the best CRM platforms for a small consulting company?” One study might open the live consumer product in a controlled browser, enter the question and wait for the answer to render. Another might observe a consenting user asking a similar question naturally.
Both observations can be useful. They do not answer the same question.
A controlled browser can make the consumer interface repeatable enough for longitudinal testing while preserving retrieval, citations, interactive modules, account state, geography, conversation history and other product-specific behaviour. A naturalistic user session adds another dimension by showing what someone actually chose to ask and how that person responded to the answer.
The method should follow the decision. A controlled browser experiment is useful when a brand wants to compare change over time under stable conditions. Naturalistic observation is useful when the brand wants to understand what people actually do. The difference is between a designed product observation and an observed human experience.
<h2>The five layers behind AI Search measurement</h2>
The measurement chain can be understood in five layers.
The first layer is <strong>information-need evidence</strong>. This is the available evidence about the questions, tasks and decisions that may inform the prompt universe. It may come from actual AI conversations collected with appropriate consent, search behaviour, site search, support logs, customer interviews, communities, surveys or third-party datasets. None of these sources captures the entire market. Each reveals a different part of it.
The second layer is the <strong>prompt set, prompt taxonomy and measurement frame</strong>. The prompt set is the collection of related prompts being measured. The prompt taxonomy classifies those prompts by dimensions such as Brand reference, Purpose, Topic, Audience and Geography. The measurement frame defines the conditions needed to interpret and compare the result.[19]
The third layer is the <strong>observation method and environment</strong>. The public evidence reviewed for this article principally concerns controlled browsers interacting with consumer products and actual user sessions observed with appropriate privacy and consent safeguards. The environment includes the AI system, model, interface or mode, Geography, language and user context where known.[19] A model endpoint can also support a separate controlled experiment, but AEO Updates did not find evidence that API-based visibility tracking is a common AEO industry practice.
The framework maps technically possible observation methods. It does not claim that every method is widely used, or that the methods appear in equal proportions across current AEO platforms.
The fourth layer is <strong>raw observations and observable evidence</strong>. A well-instrumented system may preserve the prompt, conversation context, full response, entities, response claims, associations, citations, cited sources, screenshot, browser replay, environment, timestamp and a unique observation identifier. Citations and cited sources create observable evidence relationships; they do not expose every source or retrieval process an AI system may have used.[19]
The fifth layer is <strong>core AI Search measurements and analysis</strong>. Defined calculations transform observations into measures such as Visibility, Share of Visibility, Citation Rate, Cited Source Frequency, Response Claim Frequency and Association Frequency. Measurement dimensions such as AI System, Model, Purpose, Topic, Audience, Geography, Time and Position determine how those results are segmented, while trends and benchmarking support analysis.[19]
The number on the dashboard is therefore the final layer, not the starting point. Its meaning depends on the measurement frame and observations beneath it.
<h2>There is no perfect prompt universe</h2>
AI-search demand is still difficult to observe directly. Search-volume datasets were built around conventional search boxes. First-party data captures only the customers and channels a company already sees. Shared conversation datasets are partial and can introduce privacy, sampling and demographic constraints. Strategic questions may be commercially important before meaningful demand data exists.
That makes judgement unavoidable. It does not make the resulting measurement invalid.
A prompt set may include prompts grounded in observed information needs, modelled expansions, strategic brand questions, synthetic controls or designed follow-ups. These describe how prompts were selected or constructed; they are not formal Prompt Taxonomy classes in the Starter Guide. Once selected, the prompts can be classified consistently by Brand reference, Purpose, Topic, Audience and Geography.[19]
The relevant question is not whether every prompt can be proven to represent market prevalence. It is whether the prompt set, Prompt Taxonomy and complete measurement frame are appropriate for the purpose of the programme and whether the result is interpreted within that scope.
This connects directly to the <a href="/articles/ai-visibility-no-common-denominator">denominator problem in AI visibility</a>. A score cannot be separated from the universe of questions, platforms, locations and moments over which it was calculated. The denominator defines the decision space.
<h2>Cloud browsers turn the consumer interface into a measurement instrument</h2>
Browserbase, Kernel and Steel are not AEO platforms. They are examples of infrastructure available to software that needs to operate browser sessions remotely and repeatedly.
Browserbase documents cloud sessions that can be controlled through Stagehand, Playwright, Puppeteer or Selenium. Its session tools include live viewing, network and console inspection, video recording and replay.[3] In June 2026, the company said the same infrastructure was supporting more than 35 million browser sessions each month.[4]
Kernel documents cloud-hosted browsers that expose a Chrome DevTools Protocol connection, support parallel execution, preserve state through profiles and provide live viewing.[5] Its replay system can start and stop recordings during an active session and export them as MP4 files.[6]
Steel likewise exposes remote Chrome sessions that can be controlled through established automation frameworks. Its documentation describes live viewing, session recording, persistent profiles, managed or customer-supplied proxies and geographic targeting at country, US-state and major-city levels.[7] [8]
These services show what is technically possible. A measurement system can create a browser, define location and state, open an answer engine, enter a prompt, capture the response and citations, store a screenshot or replay, timestamp the event and repeat the sequence across prompts and markets.
That is not evidence that every AEO platform uses one of these providers. A company may build comparable browser infrastructure internally, use another vendor or choose different methods for different engines. The architecture explains how this class of measurement can be performed, not which undisclosed stack a particular platform has selected.
<h2>A real browser is not a real user, and that is not a criticism</h2>
If software launches Chromium, opens an AI product and submits a prompt, the resulting answer may come from the live consumer interface. If the measurement company selected the question and automated the session, however, no consumer necessarily initiated that interaction.
The result is a controlled synthetic observation through a real consumer interface. That can be extremely valuable. It can be repeated, compared across time and inspected under stable conditions. It simply should not be interpreted as direct evidence of naturally occurring behaviour.
The same distinction exists throughout digital measurement. Synthetic monitoring uses scripted activity in a controlled environment with variables such as geography, network, device, browser and cache status set in advance. Real-user monitoring captures what actual people experience under naturally varying conditions. Synthetic measurement is repeatable and effective for detecting change. Real-user measurement is more representative of actual behaviour but harder to control.[9]
AEO needs both concepts. Controlled observations are well suited to comparing AI systems, prompt sets, Geographies and Time periods. Information-need evidence is needed to understand whether the questions being tracked resemble what people actually ask. Neither method invalidates the other because they serve different analytical purposes.
<h2>Prompt-set design and observation method are separate</h2>
The phrase “real prompts” answers one question. The phrase “real interface” answers another.
Prompt-set design describes why a question entered the prompt universe and how it is classified. Observation method describes how the response was collected. An observed human question replayed through a controlled browser combines information-need evidence with designed execution. A strategic brand prompt entered through the same consumer interface combines a chosen commercial question with live-interface observation. A naturally occurring conversation combines information-need evidence with naturalistic observation.
Every combination can produce useful information. The combination determines what can reasonably be concluded.
Google Search Console demonstrates the mirror image of the problem. Google treats an AI Mode follow-up as a new query, but publishers still cannot reliably connect a conversational-looking query fragment to the full preceding exchange and the cited response.[2] The data may reflect real behaviour while providing an incomplete environment and conversation record. A controlled browser can preserve those conditions while still relying on a designed prompt set. Each closes a different part of the measurement gap.
<h2>A worked example: one CRM prompt, two observation methods</h2>
Return to the question, “What are the best CRM platforms for a small consulting company?”
A **controlled-browser study** can enter the same prompt into the live interface, capture the rendered answer and citations, and test the effect of location or session state. It measures the product experience for a designed question, not the prevalence of that question among real buyers.
A **naturalistic user study** can observe what consenting users actually ask, how they refine the question and which recommendations influence their next action. It measures real behaviour within the sample, but it is harder to standardise and repeat.
The right method depends on the decision. A brand monitoring weekly competitive movement may prefer controlled repetition. A team exploring buyer language may need observed conversations. A product team studying citation modules may need the consumer browser. A mature programme may use both approaches and connect their findings.
<h2>Geography and user context become measurement conditions</h2>
Geography is one example of why execution context matters. Profound’s Summer 2026 Index reported that AI visibility varied materially by market. In 24 of 30 multi-region industries, at least one European market had a leading brand that did not appear in the US top five. Profound also reported a higher share of country-specific citations in markets led by regional brands.[10] The findings come from Profound’s proprietary dataset and definitions, so they should not be treated as a census of all AI-search activity. They nevertheless show why a single-market observation cannot automatically stand in for global visibility.
Cloud-browser infrastructure can make geography a controlled condition. It does not guarantee that a proxy reproduces every characteristic of a local consumer. The value lies in being able to hold location constant or compare locations deliberately.
Session state creates a similar choice. Browserbase Contexts, Kernel profiles and Steel profiles can preserve cookies, local storage, authentication and other browser data across sessions.[11] [12] [13] A measurement system can therefore compare a clean session with a persistent one, or continue a multi-turn conversation instead of restarting after every prompt.
These are design options, not universal requirements. A clean session may be the right choice for a stable benchmark. A persistent session may be more appropriate when the research question concerns continuity or personalisation. The important point is to match the condition to the decision and interpret the outcome accordingly.
<h2>Browser observation can preserve raw observations</h2>
Cloud-browser infrastructure also makes a stronger record of raw observations and observable evidence possible.
Browserbase documents video, session metadata, page and protocol events, network activity, console logs and programmatic session logs.[14] Kernel provides downloadable MP4 recordings.[6] Steel records sessions for later playback and supports current headful video replay.[15]
An AEO platform could use those capabilities to preserve the observations behind a Visibility measurement: the prompt, timestamp, AI system, Model, Geography, user context, rendered response, citations, cited sources, screenshot, replay and calculation output. A client could move backwards from a weekly result to the observations that produced it.
The infrastructure does not create the analytical chain automatically. Prompt-set design, Prompt Taxonomy, the Measurement Frame, parsing rules, measure definitions and calculation logic still have to be designed by the measurement provider. The browser contributes a replayable record of what the interface displayed.
Auditability is valuable because a metric can move even when the underlying brand has not changed. The model, retrieval system, product interface, prompt set or scoring rules may have changed instead. Raw observations make those explanations easier to investigate.
<h2>A practical Measurement Declaration</h2>
Readers do not need every technical detail before using an AEO report. They do need enough context to understand what it represents. The AEO Starter Guide calls this compact record a <strong>Measurement Declaration</strong>. It makes the denominator visible without turning methodology into a pass-or-fail test.[19]
<strong>Measure name.</strong> Is the result Visibility, Share of Visibility, Citation Rate, Cited Source Frequency, another explicit measure or a platform-native metric?
<strong>Numerator and denominator.</strong> Which successful observations form the numerator, and what full eligible observation set forms the denominator?
<strong>Unit of analysis.</strong> Is the unit a response, citation-bearing response, citation occurrence, brand mention or source?
<strong>Prompt universe.</strong> What selection method, intent-family structure, prompt count and weighting define the prompt set?
<strong>Competitive set.</strong> Which brands or category definition are used for comparative calculations?
<strong>Environment.</strong> Which AI system, Model, interface or mode, Geography, language and user context were represented?
<strong>Collection and aggregation.</strong> What cadence, repeat runs, collection volume, reporting window and failed or missing runs sit behind the result?
<strong>Platform-native metric.</strong> What is the vendor’s metric label and stated calculation method, and is the result directly observed, inferred, estimated or proprietary?
<strong>Measurement version.</strong> Which prompt-set, taxonomy, AI system, Model, parser or methodology version produced the result?
<strong>Supplemental interpretation boundary.</strong> Which conclusions does the programme support, and which questions remain outside its design?
The declaration does not rank one method above another. It makes the measurement frame, denominator and relationship between method and meaning explicit.
<h2>What public disclosure currently establishes</h2>
Profound states on an official comparison page that it runs prompts through the front-end browser interface of each answer engine daily rather than relying on API calls.[16] That is a useful high-level disclosure. It does not identify the underlying browser vendor, automation framework, account state, geography, session design or complete collection architecture.
AthenaHQ publicly describes outputs including visibility, mentions, citations, sentiment, sources and prompt-level analysis.[17] Searchable similarly describes monitoring brand appearances across major AI platforms.[18] In the reviewed public materials, AEO Updates did not find enough evidence to connect AthenaHQ or Searchable to Browserbase, Kernel, Steel or another specific browser-execution stack.
Across the named providers reviewed for this article, AEO Updates did not find evidence that API-based visibility tracking is a common disclosed practice. A model API remains a technically possible way to run a separate experiment, but the article does not use that possibility as a description of current AEO market adoption.
The absence of public detail is not evidence that a platform does or does not use browsers. It simply limits what can be stated about its implementation. The article’s purpose is not to grade providers on disclosure. It is to give readers a clearer language for understanding what different forms of measurement can represent.
<h2>The strongest systems will combine layers</h2>
No single method solves the whole problem. Information-need evidence can show what people ask but may lack complete response context. Controlled browser testing can show how a prompt set behaves over time but cannot prove that consumers naturally ask those questions. Naturalistic observation captures actual experience but introduces privacy, consent, sampling and reproducibility challenges. Business outcomes add another layer by asking whether Visibility affected awareness, consideration, traffic, leads or sales.
The strongest AEO measurement systems will connect these layers without collapsing them into one claim. They will use information-need evidence to inform the prompt universe, strategic judgement to build a useful prompt set, a stable measurement frame to make comparisons meaningful, raw observations and observable evidence to support investigation, and business outcomes to determine whether Visibility mattered.
<h2>AEO Updates Takeaway</h2>
The browser is becoming an important measurement instrument for AI search, but it is only one layer of the system. A real prompt does not necessarily mean a real user session. A real browser does not necessarily mean a real consumer. A model endpoint can support a separate experiment, but this review did not establish it as a common basis for AEO visibility tracking. The distinction is about what the evidence represents, not whether a method is inherently good or bad.
The purpose of the programme should determine the prompt set, Prompt Taxonomy, Measurement Frame, observation method and level of control. The results should then be interpreted within that design.
The strongest AEO measurement is not the system that claims to eliminate judgement. It is the system that makes its purpose, assumptions and evidence clear enough for people to use the results responsibly.
[1] <a href="https://developers.google.com/search/blog/2026/06/gen-ai-performance-reports" target="_blank" rel="noopener noreferrer">Google Search Central: “Introducing Search Generative AI performance reports in Search Console”</a>, June 3, 2026.
[2] <a href="https://support.google.com/webmasters/answer/7042828?hl=en" target="_blank" rel="noopener noreferrer">Google Search Console Help: “What are impressions, position, and clicks?”</a>, accessed August 24, 2026.
[3] <a href="https://docs.browserbase.com/platform/browser/getting-started/using-browser-session" target="_blank" rel="noopener noreferrer">Browserbase: “Using a browser session”</a>, accessed August 24, 2026.
[4] <a href="https://browserbase.com/blog/introducing-browserbase-agents/" target="_blank" rel="noopener noreferrer">Browserbase: “Introducing Browserbase Agents”</a>, June 30, 2026.
[5] <a href="https://www.kernel.sh/docs/integrations/browser-use" target="_blank" rel="noopener noreferrer">Kernel: Browser Use integration documentation</a>, accessed August 24, 2026.
[6] <a href="https://kernel.sh/docs/browsers/replays" target="_blank" rel="noopener noreferrer">Kernel: “Replays”</a>, accessed August 24, 2026.
[7] <a href="https://docs.steel.dev/overview/guides/playwright-node" target="_blank" rel="noopener noreferrer">Steel: “Automate a cloud browser with Playwright”</a>, accessed August 24, 2026.
[8] <a href="https://docs.steel.dev/overview/stealth/proxies" target="_blank" rel="noopener noreferrer">Steel: “Proxies & Proxy Rotation for Browser Automation”</a>, accessed August 24, 2026.
[9] <a href="https://developer.mozilla.org/en-US/docs/Web/Performance/Guides/Rum-vs-Synthetic" target="_blank" rel="noopener noreferrer">MDN Web Docs: “Performance Monitoring: RUM vs. synthetic monitoring”</a>, last modified October 9, 2025.
[10] <a href="https://www.tryprofound.com/reports-guides/profound-index-report-summer-2026" target="_blank" rel="noopener noreferrer">Profound: “Profound Index Report: Summer 2026”</a>, August 19, 2026.
[11] <a href="https://docs.browserbase.com/platform/browser/core-features/contexts" target="_blank" rel="noopener noreferrer">Browserbase: “Contexts”</a>, accessed August 24, 2026.
[12] <a href="https://www.kernel.sh/docs/browsers/pools" target="_blank" rel="noopener noreferrer">Kernel: “Browser Pools”</a>, accessed August 24, 2026.
[13] <a href="https://docs.steel.dev/overview/profiles-api/overview" target="_blank" rel="noopener noreferrer">Steel: “Profiles API overview”</a>, accessed August 24, 2026.
[14] <a href="https://docs.browserbase.com/platform/browser/observability/observability" target="_blank" rel="noopener noreferrer">Browserbase: “Observability”</a>, accessed August 24, 2026.
[15] <a href="https://docs.steel.dev/overview/sessions-api/embed-sessions/past-sessions" target="_blank" rel="noopener noreferrer">Steel: “Past Sessions”</a>, accessed August 24, 2026.
[16] <a href="https://www.tryprofound.com/articles/profound-vs-athenahq" target="_blank" rel="noopener noreferrer">Profound: “Profound vs. AthenaHQ: Which AI visibility platform is right for your brand?”</a>, April 16, 2026.
[17] <a href="https://athenahq.ai/platform" target="_blank" rel="noopener noreferrer">AthenaHQ platform overview</a>, accessed August 24, 2026.
[18] <a href="https://www.searchable.com/blog/geo-vs-seo-vs-aeo" target="_blank" rel="noopener noreferrer">Searchable: “GEO vs SEO vs AEO: The Honest Guide to What's Real, What's Hype, and What to Do About It”</a>, March 31, 2026.
[19] The Prompt Group: <em>The AEO Starter Guide: A Framework for AI Search Intelligence and Answer Engine Optimization</em>, Version 1.3, 2026.