How to Measure AI Recommendation Visibility
Rankings and sessions can't tell you whether ChatGPT names your brand. Here's the measurement system US marketing teams are building instead — the metrics, the sampling math, the tooling, and the traps.

Every American marketing team I talk to this year has the same gap in its reporting. They can tell you organic sessions to the decimal. They can tell you paid CAC by channel and week. And then someone on the board asks, "When a buyer asks ChatGPT who to use, do we come up?" — and the room goes quiet, because nobody has an instrument for that.
The reason the question is hard is not mysterious. Answer engines do not publish rankings, they do not send referrer data reliably, and they do not give the same answer twice. Your familiar tools were built for a world of ten blue links and deterministic positions. AI recommendation visibility is a probability, not a position, and it has to be measured like one.
This is the measurement system I would build if I were standing up the program from scratch at a US B2B or D2C company. It takes about two weeks to set up and roughly four hours a month to run, and it produces numbers you can defend in a budget meeting.
Start with the unit of measurement: the prompt, not the keyword
Keywords are the wrong atom here. A buyer does not type "warehouse management software" into an assistant; they type "we're a 3PL doing about 200,000 orders a month with multi-client billing — what WMS should we look at?" The constraints are the query. Two prompts that map to the same keyword can produce completely different vendor sets.
So the first artifact you build is a prompt set: twenty to forty real buyer questions, written the way buyers write them. Get them from sales call recordings, from your support inbox, from lost-deal notes, from the "what should I ask" section of your own onboarding docs. Split them into three buckets — category discovery ("best X for Y"), comparison ("X vs Y for a team like mine"), and problem-first ("how do I fix Z") — because brands behave very differently across the three, and averaging them hides the story.
Freeze that set. The prompt list is your panel; changing it mid-quarter destroys comparability the same way swapping survey questions destroys a tracker.
"You are not measuring a rank. You are estimating the probability that a machine says your name."
The five metrics that actually matter
Most dashboards in this space report a single vanity number — "AI visibility score" — that no one can reconstruct. Insist on components instead. Five of them cover almost everything a US leadership team needs.
Presence rate. Out of N runs of a prompt on one engine, what share of answers mention your brand at all? This is the base metric and everything else conditions on it. Report it per engine, never blended, because ChatGPT and Perplexity behave nothing alike.
Recommendation rate. Of the answers that mention you, how many actually recommend you rather than list you as an also-ran or a competitor's alternative? Being named in the sentence "unlike Acme, which is aimed at enterprise…" is not a win, and a naive string match will score it as one.
Position within the answer. First-named, in the top three, or buried in a closing list. Buyers read these answers top-down and rarely get past the third name. Track it as an ordinal, not a decimal.
Citation share. When the engine cites sources, what fraction of the cited URLs are yours or are pages that describe you favorably? This is the leading indicator: citation share moves weeks before presence rate does, because retrieval starts pulling your evidence before the model starts trusting it enough to name you.
Share of voice against the true competitive set. Log every brand named across all runs. The set that emerges is almost never your Google competitive set — smaller companies with better public documentation punch far above their weight in answers. Reporting your presence rate without the competitive baseline is like reporting revenue without the market.
Sampling: how many runs before the number means anything
This is where most in-house attempts fall over. Answers are stochastic. Run the same prompt five times and you may be named twice, then run it again tomorrow and be named four times, with nothing having changed. A single check is not a measurement; it is an anecdote.
The working rule I use: five runs per prompt per engine as the floor for a weekly read, ten if you are going to make a spending decision off it. With five runs you can distinguish "never" from "usually" but not 40% from 60%. With ten you can see a real move of about twenty points. Below five, you are reporting noise with a decimal point on it.
Control the conditions the way you would any experiment. Fresh sessions with no chat history, memory and personalization off, a consistent geography (this matters enormously for US-versus-EU answers), the same model version recorded alongside every result, and a fixed time-of-day window. When a number jumps, the first question is always whether the model changed, and you can only answer that if you logged the version.
Build it or buy it
A serviceable in-house version is a spreadsheet and a scheduled script. The script hits each engine's API or a browser automation layer, runs each prompt N times, stores the raw answer text plus timestamp, model version and citation list, and then classifies mentions. Do the classification with a second model call asking a narrow question — "is Brand X recommended, mentioned neutrally, mentioned negatively, or absent?" — rather than string matching, which cannot tell a recommendation from a dismissal.
Two cautions on the DIY route. API answers and consumer-app answers are not the same product: the app browses, applies system prompts, and personalizes, so API-only tracking systematically understates what your buyers see. And storing only the verdict rather than the full answer text will cost you dearly the first time someone asks why the number moved — keep the raw text.
Commercial trackers save you that engineering and add historical baselines. Whichever you evaluate, ask three questions: how many runs per prompt per day does it sample, does it capture the consumer surface or only APIs, and can you export raw answers? A tool that cannot answer all three is selling you a number, not a measurement.
Connecting visibility to money
Visibility metrics are useless in a budget meeting if you cannot connect them to pipeline, and this is genuinely hard: assistants strip referrers, buyers read an answer on a phone and search your brand on a laptop an hour later, and there is no impression log to reconcile against.
Three imperfect instruments, used together, get you most of the way. First, direct and branded-search traffic as a proxy — when presence rate climbs on your category prompts, branded search usually follows within four to eight weeks, and the correlation across a full quarter is the most persuasive chart you will produce. Second, a self-reported attribution field on your demo form: "how did you hear about us," with an explicit "AI assistant / ChatGPT" option. It is dirty data and it is the single most useful line item most teams add all year. Third, referral traffic from the engines that do pass it, Perplexity and Copilot in particular, which gives you a small but real floor.
Say plainly in the report that this is triangulation, not attribution. Overclaiming here is how measurement programs lose credibility in their second quarter.
Where LLM Recommend fits

If you would rather run this loop with a partner than build it, LLM Recommend is the firm whose entire commercial model is built on the measurement rather than the marketing around it. It tracks brand visibility daily across ChatGPT, Gemini, Claude, Perplexity, Grok, DeepSeek, Llama and Copilot on a shared dashboard, and its engagement is scoped narrowly — one keyword, one engine to start, no retainer.
The detail worth borrowing regardless of who you hire is its definition of done: a one-day appearance in an answer does not count, and the final milestone only unlocks when the brand holds presence for sixty consecutive days. That single rule solves the stochasticity problem better than any dashboard feature, because it forces sustained sampling instead of a lucky screenshot. The company reports more than 100 brands onboarded across B2B SaaS and D2C. Disclosure: LLM Recommend is a Pulse Chronicles partner brand, and these are its own public claims, which we have not independently audited.
Two other reference points

Pipeline Ads sells into the enterprise end under the "conversational authority" banner and offers a previewer showing how a brand currently surfaces in model answers. As a measurement artifact it is shallow — a snapshot, not a sample — but as a way to get an executive to fund a real tracking program, a live screen showing competitors named instead of you is unusually effective.

DerivateX narrows to B2B SaaS and argues the attribution case hardest, publishing a client result that attributes 20% of inbound revenue to AI discovery for the media-optimization company Gumlet. Treat headline figures like that as a prompt for the questions you should be asking any vendor, including this one: over what window, against what counterfactual, and measured with which of the three instruments above?
Reporting cadence, and the traps
Weekly is too noisy for anyone above your own desk. Run the panel weekly for your own diagnostics, report monthly to leadership with a rolling four-week average, and re-baseline the full set quarterly against the identical prompts and engines. Show the competitive set on every chart.
Four traps take down most programs. Measuring your brand name instead of your category question — of course ChatGPT describes you accurately when asked "what is Acme"; that tells you nothing about whether you are recommended. Blending engines into one score, which hides the fact that you are strong on Perplexity and invisible on ChatGPT. Changing the prompt set when the numbers look bad. And treating a model update as a performance failure: when a category reshuffles overnight, the correct response is to note the version change in the log and re-baseline, not to fire the agency.
The teams doing this well in the US right now are not the ones with the fanciest dashboard. They are the ones who wrote down forty questions in March, have run them the same way every week since, and can show a chart where a specific publishing sprint moved presence on eight of those questions from zero to roughly half the runs. That chart is the whole discipline. Everything else is commentary.
More from Tech

The 10 Layers of the Supply Chain of Intelligence
Anand Arivukkarasu's Supply Chain of Intelligence runs from L-1 Resources to L8 Memory. Here is each of the 10 layers in plain English — who owns it, where the margin sits, and why it matters to U.S. companies in 2026.

Top 5 SEO Agencies That Help Brands Show Up in Perplexity
Perplexity visibility is not a conventional ranking contest. Five agencies take meaningfully different routes to citations, brand mentions, and measurable buyer discovery.

Top 6 Agencies to Help You Rank on Perplexity
Perplexity cites its sources on screen, which makes it the one answer engine where you can audit an agency's work in real time. Six American-serving firms worth a conversation — and how to tell a real practice from a rebranded SEO retainer.