AI visibility has no common denominator

AI visibility platforms can observe the same category and produce different scores without either being wrong. The difference begins with the denominator.

The strategic challenge in AI visibility is not simply recognising a mention. It is defining the AI decision space in which that mention is supposed to matter.

Author: Ian Ash

Published: August 19, 2026

Category: Analysis

AI visibility sounds simple. A brand either appears in an AI answer or it does not. On that basis, the calculation appears to need only a numerator, the answers that mention the brand, and a denominator, the answers that were eligible to mention it.

The numerator is usually the straightforward part. The real methodological work begins with the denominator.

That matters because several AI visibility platforms can observe the same broad category and still produce different-looking scores without either result being obviously wrong. They may be counting different prompt universes, different kinds of responses, different competitive sets and different dimensions of a mention. The question is not merely whether a brand is visible. It is visible within what defined decision space?

<h2>Similar labels can describe different measures</h2>

Searchable's public documentation makes the distinction unusually clear. Its Presence Rate is the percentage of tracked prompts in which a brand appears at all. Its Visibility Score incorporates additional signals, including mention frequency, position, context quality and sentiment. Its Share of Voice is the brand's mentions as a proportion of total category mentions. Those are related measures, but they are not interchangeable.

AthenaHQ similarly describes share of voice as a brand's share of competitor mentions across a set of relevant prompts, while its documentation separately describes response-level mention rate, ranking position and cited-source data. The point is not that one platform has the definitive denominator. It is that a reporting dashboard can legitimately contain several denominators, each answering a different question.

Profound's public methodology takes the issue back another step, to the construction of the prompt universe. In a 2025 product announcement, the company said its recommendation engine draws on real user conversation data, semantic analysis and topic-level subthemes to identify prompt gaps. Its Profound Index methodology describes keyword and semantic filtering, clustering by intent and multiple prompt variations before it calculates the proportion of responses in which brands are mentioned.

These approaches share a useful premise: counting mentions is only meaningful after the measurement system has stated what it is sampling.

<h2>One set of answers can produce three defensible numbers</h2>

Consider a controlled test of 100 relevant AI responses. A brand appears in 20 of them. Its basic presence rate is 20 per cent.

Now suppose only 70 of the 100 responses name any brand in the category. The same 20 appearances represent 28.6 per cent of brand-containing responses. That second calculation is not a replacement for the first. It answers a narrower question: when the answer engine is willing to name a company at all, how often is it naming this one?

There is a third lens. If all companies together receive 110 mentions across those answers, the brand's share of voice is 18.2 per cent. Multiple brands can be named in one answer, so total category mentions can exceed the number of responses. Again, the calculation is not in conflict with presence rate or competitive inclusion. It measures a different form of competitive presence.

The difficulty begins when all three are summarised as "visibility" and presented as though they describe the same thing. They do not.

<h2>The prompt universe does more work than the headline score</h2>

The most consequential design decision may be which questions count. A tracker can be built around 50, 100 or 1,000 prompts, but the number alone says little. The important questions are whether those prompts represent how people actually ask about the category, whether branded and discovery-oriented questions are separated, whether a few familiar themes have been over-sampled, and whether commercially consequential questions receive the same treatment as broad educational ones.

Those choices shape the denominator before the system analyses a single answer. A prompt set weighted heavily toward comparison questions will tend to produce a different competitive picture from one weighted toward problem-definition questions. A set that includes many branded prompts may be useful for monitoring reputation, but it should not be mistaken for an estimate of unprompted discovery.

This is why Profound's work on real conversation data and semantic grouping deserves attention. The useful innovation is not simply a larger prompt list. It is an effort to model the question landscape around an industry rather than assuming that a traditional keyword list adequately represents it.

<h2>Near-neighbour prompts are valuable, but they are not separate markets</h2>

People rarely phrase the same need in identical language. "What is the best pain reliever for arthritis?", "What over-the-counter medicine works for joint pain?" and "What should I take for arthritis pain?" may all express a closely related underlying intent. Testing more than one formulation is valuable because answer engines can respond differently to different wording.

However, counting every variation as a wholly independent opportunity can distort the score. If one intent family is represented by 20 formulations and another by three, the first can dominate the aggregate result because the researcher generated more prompts, not necessarily because the underlying need is more important.

A stronger measurement design treats related formulations as an intent family or semantic neighbourhood. Multiple prompts then estimate performance within that family. They should not automatically make the family more commercially significant. The distinction is especially important when an executive dashboard turns a diverse prompt set into a single percentage.

<h2>Prompt tracking is repeated sampling, not a census</h2>

Most visibility tools run a defined prompt set repeatedly and report outcomes daily, weekly or over rolling periods. Searchable, for example, documents daily visibility reporting and recommends that users focus on trends rather than individual data points. Its reporting material also describes each tracked prompt as a test across AI platforms that extracts mention, position, sentiment and, where applicable, citations.

That is useful operating data. It is not an observation of every real AI interaction occurring in the market. It is repeated sampling of a selected prompt universe under particular platform, model, geography and timing conditions.

The practical implication is straightforward. A movement from 41 per cent to 44 per cent may signal a meaningful change, but the metric alone does not prove it. Readers need the underlying prompt count, test cadence, platform mix, time window and any changes to the prompt set before they can judge whether the movement is likely to be substantive.

The mature reporting view is therefore less interested in a single day's point estimate than in a transparent sequence: a daily reading for operational awareness, a rolling seven-day view for short-term direction, and a longer rolling window for strategic interpretation. Statistical uncertainty is the logical next step where a measurement provider has enough repeated observations to calculate and explain it responsibly.

<h2>Composite scores belong on dashboards, not at the end of the diagnosis</h2>

Composite visibility scores can be useful. Senior teams need a concise way to see whether a programme appears to be gaining or losing ground. The problem arises when the composite becomes the explanation rather than the summary.

A brand that appears frequently but is usually named third has a different problem from a brand that appears rarely but is strongly recommended when it does. A brand that receives citations without being named has a different opportunity from a brand that is named favourably but is not supported by cited evidence. A single number can conceal those distinctions.

The more useful approach is to keep several dimensions visible. Presence asks whether the brand appears. Prominence asks where it appears. Recommendation asks whether the answer engine actually endorses it. Sentiment asks how the brand is framed. Citations identify which sources support the answer. Share of voice describes the brand's relative competitive presence. These dimensions can inform an executive score, but they should remain available for diagnosis.

<h2>The denominator is the strategic asset</h2>

The future of AI visibility measurement will not be decided by who builds the most sophisticated mention counter. It will be decided by who builds the most credible model of the AI decision space.

That model should retain the exact prompt, the intent family, the prompt variation, the answer engine and model, the relevant geography or audience context, the response time, the full answer, the brands mentioned, their position, the cited sources and the recommendation signals. Once those records exist, a transparent presence rate can be calculated. Competitive inclusion, share of voice, market-weighted visibility and commercially weighted visibility can then be added as distinct layers rather than compressed prematurely into one score.

The numerator tells a brand whether it appeared. The denominator defines where that appearance is meant to count. In AI search, that definition is increasingly the real source of competitive advantage.

<h2>AEO Updates Takeaway</h2>

When comparing AI visibility platforms, brands should ask for the prompt universe, the unit of analysis, the competitive set, the platform mix and the time window before comparing headline scores. A lower score from one system may not indicate weaker performance. It may indicate a broader, more commercially disciplined or differently weighted denominator. The goal is not to find a single universal visibility number. It is to build a measurement system that makes the AI decision space explicit enough to support a useful decision.

<h3>References</h3>

[1] Searchable. "AI Visibility Tracking." <a href="https://docs.searchable.com/using-searchable/visibility-tracking">docs.searchable.com</a>

[2] AthenaHQ. "What Is AI Share of Voice?" 13 July 2026. <a href="https://athenahq.ai/blog/understanding-share-voice-ai-search">athenahq.ai</a>

[3] AthenaHQ. "MCP Server." <a href="https://docs.athenahq.ai/api-reference/mcp">docs.athenahq.ai</a>

[4] Profound. "Introducing Profound's data-driven prompt recommendation engine." 15 July 2025. <a href="https://www.tryprofound.com/blog/data-driven-prompt-recommendation">tryprofound.com</a>

[5] Profound. "Introducing the Profound Index." 5 November 2025. <a href="https://www.tryprofound.com/blog/introducing-profound-index">tryprofound.com</a>

[6] Searchable. "Schedule & Share AI Visibility Reports." <a href="https://www.searchable.com/workflows/automate-reporting">searchable.com</a>

Primary sources cited

This article links directly to the primary documentation, paper, filing or original reporting used for its material claims.

  1. docs.searchable.com
  2. athenahq.ai
  3. docs.athenahq.ai
  4. tryprofound.com
  5. tryprofound.com
  6. searchable.com

Continue exploring