AI answer engines broke the metric that anchored search for two decades: your average keyword position. When an assistant answers a question, there is often no ranked list to occupy and frequently no click to count. Measuring AI visibility means adopting new units built for a probabilistic, retrieval-driven system, and being honest about the real limits, that there is rarely an official feed of what users see, so credible measurement is sampling, not ground truth. This is how to build a measurement habit that survives scrutiny.
Measure four things, and keep them separate
AI visibility is not one number. Presence is whether the answer names your brand at all. Citation is whether your URL appears as a linked source. Sentiment is how you are described when you appear, from recommended option to cautionary aside. Share of voice is how often you appear relative to named competitors across the same questions. These appear to be handled by different parts of the answer pipeline and tend to move independently, so an engine can cite your data in a footnote while naming a competitor in the recommendation.
Treat it as probabilistic, so measure rates not positions
A search results page is deterministic enough to sample once. A generated answer is sampled from a probability distribution, and whether live retrieval fires can change run to run, so the same prompt can return different brands and citations across sessions. The consequence is that a single check tells you almost nothing. The meaningful unit is a rate across many samples of the same prompt over time, which is why one-off screenshots are anecdotes, not measurement.
There is no official feed, so sampling and prompt sets are the method
For most consumer AI answers there is no public API that returns exactly what a user sees, and where developer APIs exist they can differ from the consumer product in model, retrieval, and personalization. Credible measurement is therefore built on a fixed, documented prompt set, the real questions buyers ask, run on a regular cadence. Keep the set stable enough to compare month over month and large enough that one lucky run does not swing the total. Changing prompts constantly destroys the trend you are trying to see.
Web-grounded and trained-memory answers are not the same reading
When browsing or retrieval is active, the answer reflects what the engine fetched live, so fresh, crawlable content can move it quickly. When the engine answers from training data alone, you are seeing a slower reflection of how your brand was described across the web at training time. These two modes can disagree for the same brand, and the levers that change each one differ, so note where you can whether an answer was grounded before you interpret a result.
Attribution is indirect, so triangulate downstream
Much AI-answer influence leaves little or no referral trail: an in-answer mention with no click leaves nothing, and some surfaces such as Google's AI Mode strip the referrer, though some web-grounded engines (Perplexity on a citation click, AI Overviews as normal Google organic) do pass standard referrers, so standard analytics undercounts AI influence rather than missing all of it. A mention can shape a buyer's shortlist and produce nothing in your traffic report until they later search your name. The practical fix is triangulation: pair direct AI-visibility sampling with branded search volume, direct traffic, and how-did-you-hear-about-us data. None is proof alone, but together they show whether AI presence is turning into demand, and they guard against over-attributing a good month to your last content change.
- Separate presence, citation, sentiment, and share of voice: they are decided by different stages and can move in opposite directions, so a single blended number hides the real story.
- Measure rates across many samples, not positions. Generated answers are probabilistic, so one screenshot is an anecdote and a per-prompt hit rate over time is a metric.
- There is rarely an official feed of what real users see, so build measurement on a fixed, versioned prompt set of real buyer questions run on a steady cadence.
- Distinguish web-grounded answers from trained-memory answers where you can, because they can disagree for the same brand and respond to different levers.
- Much AI influence leaves little or no referral trail and standard analytics undercounts it (though some web-grounded engines do pass referrers), so triangulate sampling with branded search and direct-traffic signals, and judge progress over a quarter, not a single noisy month.
