The Report Said 61 Percent. Nobody Could Say 61 Percent of What.
A CMO sent me a report last month. Headline number: 61 percent "AI visibility." Her board wanted to know if that was good. I asked a different question: 61 percent of what? Neither of us could point to a single sentence a real buyer had typed to produce it.
The vendor couldn't answer cleanly. The audit ran twenty prompts against major answer engines and counted how often the brand showed up. The twenty prompts were built from category keywords — "best project management software," "top CRM for small business," the kind of terms an SEO tool would suggest. Nobody had asked the sales team what prospects actually type into a search box, or say to a rep, before they buy.
We rebuilt the question set from ninety days of real sales-call transcripts and support-ticket language. Same brand, same answer engines, same measurement window. The defensible baseline came in at 22 percent. Not 61. A 39-point gap, and the report the board was about to act on was measuring a market that didn't exist.
That gap is not a story about a bad vendor. It's a story about a category-wide mistake. Most AI-visibility numbers in circulation today are answers to a question nobody wrote down carefully enough to defend.
The Prompt Set Is the Instrument, Not the Score
Treat an AI-visibility score the way a researcher treats a survey result: it is only as good as the sampling method behind it. A political poll of twenty people from one zip code is not "wrong" arithmetic — the math is fine. It is wrong because the sample was never the population it claims to represent.
The same failure sits inside most AI-visibility audits, and it is easy to miss because the output looks rigorous. A percentage, a chart, a trend line. What's invisible is the list of prompts that produced it, and whether that list was ever checked against how real buyers actually talk. Most vendor audits never publish the prompt list at all, so nobody outside the vendor can check whether the sample matched the market it claimed to describe.
Twenty generic category prompts will always return a number. They will almost never return a number anyone should reallocate a marketing budget against, because they measure the vendor's guess about buyer language, not the language itself. In my work, I've come to treat the prompt set as the actual instrument being tested — the score is just a readout. Fix the instrument first. Everything downstream inherits its quality, or its flaws.
A Five-Step Method for Building a Defensible Question Set
When I rebuild a question set for a client, I follow the same five steps regardless of industry.
- Source real buyer language, not guessed keywords. Pull from support transcripts, sales-call notes, site search queries, and search console data — the actual words prospects use, not the words a content team assumes they use.
- Segment by purchase stage. Awareness, comparison, objection-handling, and post-purchase questions behave differently inside answer engines, and lumping them together hides which stage is actually underrepresented.
- Include competitor-comparison and "vs" phrasing explicitly. Category terms alone miss the single most common way buyers actually evaluate — by name, against a named alternative.
- Include reputational and failure-mode questions. "Is X reliable," "is X worth it," "did X get acquired" — brands routinely leave these out, and they carry disproportionate weight in how answer engines characterize a company.
- Set a refresh cadence, quarterly at minimum. Buyer language shifts faster than SEO keyword lists ever did, and a question set is a snapshot, not a fixture.
None of these steps require special tooling. They require discipline about where the questions come from, and a willingness to throw out prompts that sound plausible but were never actually asked.
What the Real Baseline Changes
The honest number is usually lower than the vendor number, and that's the point. A lower, defensible baseline doesn't just correct a metric — it changes where you spend. That's uncomfortable in a board meeting, and it's exactly why most vendors don't build their prompt sets this way in the first place.
In the case above, once we rebuilt the question set, close to 40 percent of the real buyer queries fell into comparison phrasing ("X vs Y for a fifty-person team") and reputational phrasing ("is X still around"). The original audit had zero questions in either category. It wasn't measuring the market wrong by a little. It was measuring a different market.
That reallocation conversation is where the exercise earns its keep. Two case-study pages already dominated the real question set — coverage there was already strong, and more investment in that direction would have been wasted. The gap was an entire comparison-page category that simply didn't exist. Once it was built, coverage on that segment moved within a single content cycle.
I won't attach a precise after-number to that here — real movement in answer engines is directional and slower than any dashboard implies, and a suspiciously clean before-and-after is usually a sign the measurement is still broken somewhere. What I can say is the direction was unmistakable, and it was pointed at the exact gap the rebuilt question set had exposed.
That's the actual value of doing this work properly. Not a better score. A correct map of where the content debt actually sits.
The Trap Is Treating the Baseline as Fixed
Buyer language moves. A question set built once and left alone starts drifting the moment it's finished. I've seen sets older than two quarters where 15 to 20 percent of the questions no longer matched how the sales team said prospects were actually asking — new competitors entered the comparison language, a feature the market used to ask about stopped mattering, a new objection replaced an old one.
A static baseline doesn't fail loudly. It fails quietly, by continuing to produce a number that looks stable while the market underneath it has already moved. Treat the question set the way you'd treat any research instrument that measures a moving target: revisit it on a fixed schedule, not when someone happens to remember.
The score is not the deliverable. The question set is. Get that instrument right, and the score — whatever it turns out to be — is finally worth trusting enough to act on. I write more on GEO and AI-search measurement on my blog.
Written by Brian, Dr. Jonah Tebaa's AI partner, on his behalf.
