The dashboard says your brand visibility in AI answers rose 23 percent. Nice. Before anyone moves budget, ask what the score actually measured.
AI visibility tools can watch whether ChatGPT, Gemini, Copilot, Perplexity, and other answer engines mention a company. They can also disagree with one another while looking equally serious in a slide deck.
On August 3, the Interactive Advertising Bureau published Measuring Visibility in the AI Era, a 37-page framework for evaluating this young category. IAB says more than 20 companies now sell AI visibility measurement, often with different prompt libraries, platform coverage, and scoring rules.
That gives marketing leaders a useful standard before procurement gets carried away.
A precise-looking score can still be a weather report wearing a necktie.
Start with the decision
The framework draws a clean line between directional and decision-grade measurement. Directional data can flag a pattern, suggest an issue, or support an internal briefing. It is not strong enough by itself for budget allocation, agency reviews, or executive strategy.
Decision-grade data needs substantially more rigor across sample size, query volume, prompt types, collection cadence, reproducibility, validation, documentation, and platform coverage. The full IAB framework spells out those distinctions and the disclosures providers should make.
So write down the decision first. Are you watching for a possible reputation problem? Comparing two content themes? Cutting a six-figure media or agency budget? Those jobs do not deserve the same evidence threshold.
Ask these seven questions
1. Which prompts are in the sample?
A visibility score is sensitive to the questions a provider chooses. A Milwaukee law firm can look dominant in a prompt set built around broad legal definitions and invisible in one built around real hiring questions.
Ask where the prompts came from, how many are used, which audience or buying stage they represent, and how often the library changes. Ask to see examples. A provider may protect some proprietary details, but “trust our methodology” is not a methodology disclosure.
2. Which platforms, models, locations, and account states are covered?
Results can change by platform, model version, geography, time, and whether a session is logged in. A blended score can hide that variation.
Require the platform and model mix behind the number. Ask whether tests run in clean sessions, how location is controlled, and whether platform updates trigger a baseline reset. If a Wisconsin business appears in one answer because the session already knows the user lives nearby, that should not be reported as universal national visibility.
3. How many times is each question run?
AI responses are not deterministic. The same prompt can name three firms on Monday and a different five on Tuesday.
IAB recommends documenting variability and reproducibility. Ask how many repeated runs support each reported result, what the normal variance looks like, and what change is large enough to count as a signal. One screenshot proves an answer happened once. It does not prove a trend.
4. What does the score count?
IAB organizes brand metrics into four layers:
- Presence: Does the brand appear or earn a citation?
- Prominence: Where does it appear, and how visible is it?
- Portrayal: Is the description accurate, useful, and framed fairly?
- Persuasion: Does the appearance support a recommendation or action?
A mention and a recommendation are different events. So are a linked citation, a named reference, and an uncited claim. Ask the provider to define each one and explain how the final score weights them.
5. Can you inspect the underlying answers?
A useful tool should let the team move from a chart to evidence. You need the prompt, response, citation, timestamp, platform, and model context for a meaningful sample of results.
This matters when the score drops. Maybe competitors gained ground. Maybe the model changed. Maybe the prompt set changed. Maybe one false claim appeared repeatedly. Without the underlying records, the team gets a red arrow and a meeting.
6. How are errors classified?
The IAB framework separates hallucinations from factual inaccuracies. A hallucination is fabricated by the AI system. A factual inaccuracy can trace back to an old or wrong source that the AI repeated faithfully.
The remedy depends on the cause. A bad address copied from an old directory calls for source cleanup. A made-up service attributed to the company calls for platform reporting, monitoring, and clear owned-source evidence. Ask whether the tool surfaces questionable answers for human review or quietly rolls them into a percentage.
7. What business action can this evidence support?
Close the loop. A directional finding might justify checking a page, updating an outdated profile, or running a deeper test. It should not automatically trigger a new content factory.
For higher-stakes decisions, connect visibility evidence with first-party outcomes: qualified visits, branded searches, consultation requests, lead quality, sales conversations, or corrected reputation issues. Our marketing causality check offers the same warning for every dashboard: a reporting model should earn the decision it is being asked to make.
A simple procurement brief
Put the seven answers into one page before buying a tool:
- the decision and stakes;
- the audience, category, geography, and prompt source;
- platforms, model versions, session state, and collection dates;
- repeat-run count and expected variance;
- metric definitions and score weights;
- access to underlying responses and citations;
- the rule for moving from observation to action.
Then label every recurring report either directional or decision-grade. That small bit of honesty can prevent a curious signal from becoming an expensive instruction.
Use the score to investigate, not decorate
AI visibility measurement is useful when it helps a team find a real problem: a missing source, weak proof, an inaccurate description, or a competitor earning stronger recommendations for a clear reason.
It gets silly when the score becomes another number to make green before the quarterly meeting.
If the evidence points to weak owned content, start with the pages and queries already showing signs of demand. If the team is tempted to produce a stack of nearly identical answers, read our warning about AI answer bait. Better measurement should make the content calendar more selective.
Does your AI visibility report support a decision, or just a slide?
We can audit the prompts, evidence, content gaps, and reporting rules before the score starts steering the strategy.
