AI visibility is how often, and how accurately, an AI assistant mentions or cites your brand when someone asks it a question. It is measured by running a fixed set of prompts through engines like ChatGPT, Perplexity and Gemini many times over, then reporting the results as a range rather than a single number, because the same question asked twice does not return the same answer.
That last part is the piece most people get wrong, and it is the reason two AI visibility tools can look at the same brand in the same week and disagree completely. This page explains what the metric actually is, what the published research says about how brands really perform, and what a measurement has to look like before you are allowed to make a decision with it.
What is AI visibility?
AI visibility is your brand's presence inside AI-generated answers. Not your position on a results page, but whether the assistant names you, cites your website as a source, and describes you accurately when a real buyer asks it something in your category.
It goes by several names, and the research literature treats them as one family. In Generative Engine Optimization at Scale, Pratyush Kumar of Ranqo writes that the shift is "variously called Generative Engine Optimization (GEO), Answer Engine Optimization (AEO), and AI Search Visibility", and states plainly: "We treat AEO and AI Visibility as part of GEO" (Kumar, arXiv, 18 June 2026). You will also see the same thing called LLM visibility, particularly when the discussion is about a specific model rather than a search product.
So the vocabulary is unsettled, but the object being measured is not. Three things are being counted:
- Mention. Did the assistant name your brand at all?
- Citation. Did it link your website as the source it drew from?
- Framing. How did it describe you, and was that description correct?
A brand can be mentioned without being cited, and cited without being mentioned favourably. Those are three separate outcomes, and collapsing them into one score is where most reporting starts to mislead.
How AI visibility differs from a search ranking
A Google ranking is a position on a page that is broadly stable and that you can check once. AI visibility is a probability distribution that changes every time you sample it.
Julius Schulte, Malte Bleeker and Philipp Kaufmann make the distinction precisely in Don't Measure Once: "In classical search engines, results are comparatively transparent and stable: a single query often provides a representative snapshot of where a page or brand appears relative to competitors. The inherent probabilistic nature of AI search changes this paradigm. Answers can vary across runs, prompts, and time, making one-off observations unreliable" (Schulte, Bleeker and Kaufmann, arXiv, 8 April 2026).
Their conclusion is the sentence that should govern every AI visibility report you ever read. Visibility must be characterised "as a distribution rather than a single-point outcome".
| Search ranking | AI visibility | |
| What is measured | Position of a URL | Whether a brand is mentioned, cited and framed correctly |
| Stability | Broadly stable between crawls | Varies between two identical runs |
| Sample needed | One check is usually representative | Tens to hundreds of repeated queries |
| Correct output | A number | A range with a confidence interval |
| Who supplies the data | Search Console, first party | Third party tools, no official source |
| Failure mode | You rank lower than you thought | You are cited for something you never said |

How likely is your brand to be mentioned at all?
Far less likely than most agencies imply, and the gap between large and small brands is the single most important number in this field.
Kumar's study analysed more than 100,000 prompt responses across more than 100 brands. The visibility rates split sharply by brand size (Kumar, arXiv, 18 June 2026).
| Brand tier | Visibility rate in AI answers |
| Global household names | 73% |
| Established mid-market and regional brands | 44% |
| Niche and small brands | 11% |

Kumar is direct about who the difficult case is: "The hard case is everyone outside the already-authoritative top brands, SMEs, D2C brands, creators, and early-stage startups."
If you are a small or mid-sized business, an 11% baseline is your realistic starting point. Any tool or consultant presenting a low number as a crisis, or a modest number as a failure, is working without a benchmark.
The AI search visibility metrics and KPIs that actually matter
There is no single AI visibility metric. The AI search visibility metrics and KPIs worth reporting are a small set of separate measurements, and a credible report keeps them separate rather than averaging them into one number.
| Metric | What it answers | How stable is it? |
| Mention rate | In what share of answers is the brand named? | Most stable of the set |
| Citation share | Of all sources cited, what share are yours? | Moderate, needs large samples |
| Share of voice | How often are you named versus named competitors? | Moderate |
| Position within the answer | Are you first in the list or eighth? | Low, order shifts between runs |
| Sentiment and framing | Are you described positively, and correctly? | Lowest, see below |
| Accuracy of attribution | Is the claim attributed to you one you actually made? | Rarely measured at all |
Sentiment deserves a warning. Kumar's data found that whether a brand is framed positively or negatively flips about 6.7 times more often than whether it is mentioned at all (Kumar, arXiv, 18 June 2026). A sentiment score taken from a single measurement round is close to meaningless.
Why two tools give the same brand two different scores
Because the tools are measuring differently, and until August 2026 there was no shared standard telling them how.
The Interactive Advertising Bureau released Measuring Visibility in the AI Era on 3 August 2026, and the reason it gives for existing is the problem itself. More than 20 companies now sell AI visibility measurement tools, each using different methodologies that can produce different answers for the same brand.
The IAB framework does not rank vendors or endorse tools. It supplies shared vocabulary, quality criteria and disclosure requirements. Its most useful contribution is a two-tier split that anyone commissioning this work should adopt immediately.
- Directional measurement. Identifies patterns and signals trends. Useful for early detection. Explicitly not sufficient for allocating budget.
- Decision-grade measurement. Meets a higher standard across six criteria: query volume, sample size, prompt type coverage, testing cadence, reproducibility and platform coverage.
Six criteria. If a report does not disclose all six, it is directional, whatever the cover says. That distinction is the most practical thing to take from this page.
How many queries does a reliable measurement need?
Between roughly 40 and 150 per platform, depending on the platform and how precise you need to be.
Ronald Sielinski of IQRush answered this quantitatively in Quantifying Uncertainty in AI Visibility. Treating citation share as a statistical estimator rather than a fixed value, and sampling repeatedly across Perplexity Search, OpenAI SearchGPT and Google Gemini, he derived the sample sizes needed for a 95% confidence interval of roughly plus or minus 2.5 percentage points on citation share (Sielinski, IQRush, arXiv, 9 June 2026).
| Platform | Queries needed for a 95% CI of about 2.5 points either way |
| Google Gemini | roughly 40 to 50 |
| Perplexity | roughly 90 to 100 |
| OpenAI SearchGPT | 150 or more |
Sielinski adds a caution for SearchGPT. Its non-monotonic convergence means a fixed commitment to the full sample is better than stopping early once the number looks settled.
The instability underneath those numbers is severe. Measuring the overlap between the sets of domains cited on repeated runs of the same query, he found a median Jaccard overlap of 0.29 to 0.31 on Gemini, 0.33 to 0.40 on SearchGPT and 0.50 on Perplexity. The rate at which two runs returned an identical citation set was between 0.01% and 0.10% on Gemini.
Ask Gemini the same question twice and you will effectively never get the same sources back. That is the reality any honest measurement has to be built on.
His worked example is the one to remember when someone shows you a competitor comparison. A domain at 9.5% citation share with a 95% confidence interval of 5.5% to 12.5%, against a competitor at 6.0% with an interval of 4.0% to 8.0%, is "statistically indistinguishable" despite an apparent 3.5 point lead.
A mention is not always real: fabricated citations
A significant share of the citations AI systems produce do not hold up when you check them, and counterintuitively the problem is worse for well-known brands.
Zoltan Varga tested 100 B2B entities across 1,400 probe runs and 2,062 sources in Per-Entity Bias Mapping for AI Visibility, and found Tier 1 brands returned 52.69% fabricated citations against 37.87% for Tier 3 entities, a difference of 14.82 percentage points at p=1.67e-11 (Varga, arXiv, 19 June 2026). Under regulatory-framed queries the rate rose to 56.77% against a 37.59% baseline, a jump of 19.2 points.
Varga names the mechanism the "Brand Hallucination Paradox". The model's familiarity with a well-known brand makes it more willing to generate a plausible but incorrect completion about it. He describes the escalation pattern as "rejection-induced confabulation escalation", where compliance filters paradoxically amplify hallucination rates for larger brands.
The practical consequence is uncomfortable. A rising mention count can coexist with rising misinformation about you. Varga's conclusion is that existing frameworks focus on citation rate and mention frequency but do not systematically distinguish verified mentions from false attributions. If your report does not verify mentions, it is counting some things that are not true.
Where AI engines actually take their citations from
Overwhelmingly from company websites, and after that from third party list articles. Kumar's citation breakdown found 78% of citations go to corporate websites and 21% are "best-of" listicles (Kumar, arXiv, 18 June 2026).
Two things follow. First, your own site remains the primary asset, which means the structure, clarity and factual consistency of your pages is doing most of the work. Second, roughly one citation in five comes from a page you do not own, on somebody else's domain, which is why presence in credible third party round-ups is not a vanity exercise.
How to measure AI visibility, step by step
A defensible measurement has six parts. Skip any of them and you have a directional signal, not a decision.
- Fix a prompt set. Write the questions a real buyer would ask, in natural language, covering informational, comparison and commercial intent. Freeze the list so results stay comparable over time.
- Name your competitor set. Share of voice is meaningless without a defined field.
- Choose platforms deliberately. ChatGPT, Perplexity and Gemini behave differently and need different sample sizes. Report each separately before combining anything.
- Repeat, at the volumes above. One run per prompt is not a measurement.
- Verify the mentions. Check that cited pages exist and that claims attributed to you are ones you made. Given Varga's figures, this step is not optional.
- Report ranges, not numbers. Every figure gets a confidence interval, and any comparison whose intervals overlap is reported as no difference. Applied consistently, that turns a list of AI search visibility metrics and KPIs into something you can actually make a decision on.
What counts as a good AI visibility score?
There is no universal threshold, and any tool presenting a score out of 100 without disclosing its methodology is presenting an opinion.
The honest benchmarks are the ones from published research rather than vendor dashboards: 11% for niche and small brands, 44% for established mid-market and regional brands, 73% for global household names (Kumar, arXiv, 18 June 2026). Judge yourself against your own tier and against your own earlier measurement, taken the same way.
Two rules make a score trustworthy, whether the vendor calls it AI visibility or LLM visibility. It should be traceable to a disclosed prompt set and sample size, and it should move slowly. A visibility score that swings wildly week to week is usually measuring its own sampling noise.
How to improve brand visibility in AI search engines
The evidence points at three levers, in order of how much of the citation pool each one controls.
Make your own site the best available source
It accounts for 78% of citations. Answer questions directly and early on the page, keep facts consistent across every page, and make sure the same claim about your business does not appear three different ways in three different places. Inconsistency is what produces a confused or fabricated description later. This is the substance of AI SEO work, and it starts with the site itself.
Earn presence in third party round-ups
They carry 21% of citations, and they are pages you do not control. Being genuinely listed in credible comparisons and directories is the practical route into that share. We set out the full method in our guide to getting your website cited by AI chatbots.
Fix your entity data before chasing mentions
If AI systems cannot reliably establish who you are, what you do and where you operate, extra mentions simply increase the chance of an inaccurate one. Consistent naming, consistent descriptions and corroboration across independent sources come before volume.
One caution. Because sentiment flips 6.7 times more often than mentions, do not chase a sentiment score. Improve the accuracy and consistency of the underlying facts and let the framing follow.