
A colleague of mine runs digital strategy for a direct-to-consumer brand doing roughly forty million in annual revenue. Last month, their CEO asked for a report on the company's AI-search visibility.
It was a perfectly reasonable request. ChatGPT, Perplexity, and Google's AI Overviews now drive a meaningful share of their discovery traffic. The CEO wanted to know three things: Are we showing up? Are we being cited accurately? Where are we losing to competitors?
My colleague spent three days trying to produce that report.
She had no dashboard. No standardized metric. No vendor tool she could buy off the shelf. The data existed — scattered across chat logs, manual spot-checks, a handful of scraping outputs from an intern's side project — but she had no coherent framework for assembling it into something her CEO could act on.
She is not alone. This is the measurement layer of AI-search visibility, and for most organizations right now, it does not exist.
The Measurement Black Hole
SEO has a mature measurement ecosystem. Search Console tells you which queries surface your pages. Ahrefs and SEMrush model your competitive positioning. Rank trackers give you daily movement data across thousands of keywords. You can walk into any boardroom and produce a credible, defensible answer to the question: "How are we doing in search?"
Generative engine optimization has none of this.
The AI models that increasingly mediate discovery — large language models powering ChatGPT, Perplexity, Claude, Gemini, and the AI-generated summaries threaded into traditional search results — do not publish query logs. They do not provide webmaster tools. They do not expose an API that tells you whether your brand was cited, in what context, or with what frequency in any given week.
The result is a measurement gap that would be unthinkable in any other channel. Imagine running paid search without conversion tracking. Imagine managing a PR program without media monitoring. That is the current state of AI-search visibility assessment. Brands know the channel matters. They cannot say with precision how well they are performing in it.
The cost of not measuring it is itself measurable. In a study of real Google browsing behaviour published in July 2025, Pew Research Center found that "Users who encountered an AI summary clicked on a traditional search result link in 8% of all visits", while those who did not see a summary clicked through nearly twice as often, on 15% of visits (Pew Research Center, July 2025). The traffic your analytics can still see is being reshaped by a layer your analytics cannot.
This gap is not a temporary inconvenience. It is a structural feature of how LLM-based search works. Unlike traditional search engines, which index a relatively stable corpus and return ranked lists of links, an AI-assisted search response is generated stochastically. The same query asked five times can produce five different answers, five different citation sets, five different brand mentions — or none at all. Measuring that requires a different conceptual model than the one we inherited from SEO.
And yet, measurement is not impossible. It simply requires us to define what we mean by visibility before we attempt to quantify it.
Three Dimensions Worth Tracking
In my work with organizations navigating the AI-search transition, I have found it useful to decompose visibility into three separate dimensions. Each answers a different question. Each requires a different measurement approach. Together, they form a coherent picture.
1. Presence
Presence is the simplest dimension: when a relevant question is asked, does your brand appear at all?
This is not the same as ranking. An AI-generated answer does not produce a ranked list; it produces a narrative. Your brand might appear as a cited source in that narrative. It might appear as an unlinked mention. It might appear through a product, a person, a data point, or a quoted passage. Presence asks whether, across a defined set of queries that matter to your business, your organization registers in the output in any form.
Measuring presence requires building a query set — the questions your customers and prospects actually ask — and sampling responses at a frequency that captures the inherent variability. Once a week is not enough when the same query can return a different answer every time. Twice a day across a representative sample is closer to useful.
Presence is binary at the query level but becomes a rate across your query set. A brand cited in 62% of responses for its priority queries has a presence rate of 62%. That number, tracked over time, is a leading indicator of discovery risk.
2. Accuracy
Accuracy is where measurement gets harder — and more consequential.
Showing up is necessary but insufficient. If an AI response cites your brand and gets the facts wrong, you have not gained visibility. You have gained a correction burden. I have seen organizations cited in AI-generated responses that attributed products they do not sell, capabilities they do not have, and pricing that was years out of date. The model confidently asserted these things, and the user had no way of knowing they were false.
Accuracy measurement requires human review. There is no automated way to reliably determine whether an AI-generated statement about your organization is factually correct, because the ground truth resides in your internal knowledge — product specifications, service parameters, leadership bios, current positioning — not in any public corpus the model was trained on.
In practice, this means sampling responses and classifying each citation as accurate, partially accurate, or inaccurate. An accuracy rate below 85% is a problem. Below 70% is a liability. Organizations that track this dimension consistently tend to discover errors they would otherwise never know existed, because their customers see the AI's answer but rarely report back to the brand about what was said.
3. Positioning
Positioning captures the competitive dimension: when your brand appears, how does it appear relative to alternatives?
An AI response might mention your organization in paragraph four while featuring a competitor in paragraph one. It might describe your offering as "a solid mid-market option" while characterizing a competitor as "the leading solution." It might cite you for a narrow sub-topic while crediting someone else for the broader category.
This is the dimension closest to what SEO practitioners think of as ranking, but the analogy is imperfect. Positioning in generative search is qualitative and narrative. It is not a number on a scale of one to ten. It is the relative prominence, framing, and recommendation weight the model assigns to your brand within a generated answer.
Tracking positioning requires scoring each response on a simple ordinal scale — primary mention, secondary mention, or tertiary mention — and tracking the distribution over time. Brands that consistently appear as a primary or secondary mention for their priority queries are building AI-search equity. Brands that appear only as tertiary mentions, or not at all, are ceding ground they may not know they are losing.
Building a Scorecard That Actually Works
These three dimensions — presence, accuracy, positioning — provide the conceptual structure. The practical challenge is implementing a measurement cadence that is rigorous enough to be useful without being so expensive that it gets abandoned after two quarters.
Below is the process I recommend to organizations building their first AI-search visibility scorecard. It is deliberately manual at the start. Automation comes later, once you know what you are measuring and why.
A Five-Step Scorecard Build
- Define your query set. Identify the thirty to fifty questions your customers and prospects actually ask — not the keywords you wish they searched for, but the real questions your sales team fields, your support team answers, and your analytics show arriving through long-tail organic search. These become your measurement universe.
- Select your AI surfaces. Choose the platforms that matter for your audience. For most B2B organizations, ChatGPT and Perplexity are the starting point. Consumer brands should include Google AI Overviews. Technology and research-heavy sectors should add Claude. Do not attempt to cover everything; cover what your customers actually use.
- Establish a sampling cadence. Query each surface with each question in your set at a frequency that captures response variability. Twice daily is a practical starting point — morning and evening — across a rolling subset of your query set so the total volume stays manageable. For a fifty-question set across three platforms, that is three hundred responses per day. A dedicated analyst or a well-configured scraping layer can handle this.
- Score each response across all three dimensions. For presence: was the brand cited in any form? For accuracy: was the citation factually correct? For positioning: was the brand a primary, secondary, or tertiary mention relative to competitors? Log the results in a simple structured format — a spreadsheet works for the first quarter.
- Review monthly at the leadership level. The scorecard is not an operational dashboard. It is a strategic instrument. Review it with the same cadence and seriousness you apply to brand health tracking or net revenue retention. Trends matter more than point-in-time numbers. A declining presence rate over three consecutive months is a signal that demands a response.
This is not a finished product. It is a starting framework — the minimum viable measurement layer for a channel that has outgrown the "we should probably pay attention to that" stage and entered the "we need to manage this systematically" stage.
This is applied AI in its least glamorous form. Not the model architecture, not the prompt engineering — but the measurement layer that tells you whether any of it is working. AI-search visibility, for all its technical novelty, depends on the same discipline: define what success looks like, measure it honestly, and act on what you learn.
A Channel Without a Dashboard
The absence of standardized measurement for AI-search visibility is not an industry failure. It is a predictable stage in the evolution of any new discovery channel. Search engines existed for years before Google launched Search Console. Social media analytics were primitive for half a decade after brands started spending real money on Facebook.
The organizations that build measurement competence now — even imperfect, manual, first-generation measurement — will have an asymmetric advantage when vendor tools eventually arrive. They will know what to measure because they have already done the conceptual work. They will evaluate tools against real operational needs, not marketing feature lists.
My colleague with the forty-million-dollar brand eventually produced her report. It took a week, not three days, and it was built on a spreadsheet, not a SaaS dashboard. But it gave her CEO something no vendor could have sold him: an honest answer about where they stood, backed by a repeatable methodology and anchored to the questions their actual customers ask.
That, for now, is what good looks like.
I work with a limited number of organizations each quarter on AI-search visibility strategy and measurement. If your team is ready to move beyond guesswork, reach out for an introductory conversation.
Written by Brian, Dr. Jonah Tebaa's AI partner, on his behalf.