A client asked us something last month that our reporting stack had no answer for.
“Why does ChatGPT recommend them and not us?”
We did what most agencies do. Opened four tabs. Typed the prompt into ChatGPT, Perplexity, Gemini and Claude. Screenshotted the two answers that made us look good. Put them on slide 11 under the heading “AI Search Observations.”
That is not measurement. That is a vibe with a timestamp.
I have been in digital marketing for twenty five years and I have watched this exact moment happen twice before. Once when organic search stopped being about keyword density and started being about links. Once when social stopped being about follower counts and started being about paid distribution. Both times, the agencies that survived were not the ones with the best opinions. They were the ones who built a way to measure the new thing first, and could show a client a number on a Tuesday.
This is the third time. Here is what I think it takes.
Manual prompting is not a methodology
Start with why the screenshot approach fails, because most agencies are still doing it and quietly hoping it holds.
Answer engines are not deterministic. Ask the same question twice and you can get two different sets of cited sources. Ask it from a different account, in a different country, with a different chat history, and the answer moves again. A screenshot captures one roll of the dice and presents it as the state of the world.
Then there is coverage. A single client might have 30 questions that genuinely matter, the ones that sit right before a purchase decision. Across 4 engines, checked weekly, that is 480 checks a month for one account. Nobody’s intern is doing that. If someone tells you they are, they are doing 8 of them and rounding up.
We got the scope wrong the first time we tried this properly. I had the team track 200 prompts on one account, on the theory that more coverage meant more rigour. The output was unreadable. Most of those prompts were questions no buyer would ever type, and the noise buried the 20 or so that mattered. We cut it to 30 and the pattern showed up in the first week. I had assumed this was a volume problem. It was a selection problem.
And then there is the part that actually matters, which is comparison over time. One screenshot tells you nothing. The only useful sentence in this entire category is “you were cited in 18 percent of the answers that matter in March, and 31 percent in June, and here is what moved it.” You cannot get to that sentence by opening tabs.
Citation is the new impression
The structural change is smaller than the hype and more brutal than it sounds.
Search used to hand the user a list and let them choose. That list had ten slots on page one, and a long tail of positions below it that still leaked some traffic. Being mediocre had a floor. You ranked eleventh, you got something.
An answer has no list. It has a paragraph, and inside that paragraph there are maybe three sources the model decided to lean on. There is no page two to lose. You are cited, or you do not exist for that question. The distribution is not a gentle curve any more, it is closer to a cliff.
This changes what a good report looks like. Position is the wrong unit. The right unit is share: out of the answers that matter to this business, how often does this brand appear, and against whom. Share of citation as share of voice. That framing is the whole shift, and once an agency adopts it, the rest of the work reorganizes itself around it.
Two numbers, not one
The mistake I see in early tooling is collapsing this into a single visibility percentage. It reads well on a dashboard and it is useless for deciding what to do on Monday.
There are two separate questions, and they fail for different reasons.
Does the model have a reason to trust this brand? This is about whether the entity is well defined, consistently described across the sources models actually retrieve from, and corroborated by third parties who are not you. Call it authority. It moves slowly and it is mostly earned off your own domain.
Does the model actually pick this brand for this question? This is about whether your content answers the specific intent, in a shape a retrieval system can lift, on a page the crawler can reach. Call it visibility. It moves faster and it is mostly fixable on your own site.
A brand can score well on the first and badly on the second, which means the story exists but the answer to that particular question is missing. A brand can score badly on the first and occasionally well on the second, which is usually a fluke and will not survive the next model update. The prescription is different in each case. One number cannot carry that.
Split by intent as well. Being cited for “what is X” and invisible for “best X for enterprise” is not a partial win. It is a total loss on the only question that had money attached to it.
Why this is an agency problem specifically
Client side teams can get away with a rough sense of this for another year. Agencies cannot, for three reasons.
You sell proof. The entire retainer model rests on being able to show that something you did caused something to move. The moment a meaningful share of buyer research happens in a surface you do not measure, your proof has a hole in it. Clients notice holes.
Your buyer is going to ask first. They are already using these tools personally. The question arrives in a meeting, unscheduled, and whoever has a real answer that day gets remembered.
It is the cleanest new line item in years. Not a rebrand of what you already sell. New work, new deliverable, measurable, and genuinely hard to do by hand. That combination is what defensible scope is made of.
What the work actually looks like
The loop is not complicated. It is just tedious enough that it has to be instrumented.
Start by mapping intents, not keywords. Write the questions a real buyer asks in the last two weeks before they choose, in the words they would use, including the comparison ones with your competitors’ names in them. Thirty is usually enough to be honest, and honest beats comprehensive.
Probe those intents across engines on a schedule, and record who got cited, not just whether you did. The competitor set that shows up in answers is frequently not the competitor set in the client’s deck. That finding alone has justified the exercise in every account we have run it on.
Find the gaps and read them properly. A gap is not “we need a blog post.” A gap is a specific question where a specific competitor is being cited from a specific source type, and the reason is usually one of four things: they have the answer and you do not, their version is structured to be liftable and yours is buried in a case study PDF, they are corroborated somewhere you are not, or your page is technically unreachable to the crawler that matters.
Then publish against the gap, wait, re-probe, and measure the delta. Probe, gap, brief, publish, re-probe, measure. That is the whole method. The value is not in any single step, it is in closing the loop so the next brief is written from evidence instead of instinct.
We built our own instrument for this because in early 2026 there was nothing that modelled citation per intent and per engine the way we needed, and I got tired of my team opening tabs. That became AVO. Whether you use ours, someone else’s, or something you build in a weekend matters far less than whether you are running the loop at all.
The part that should worry you
Ranking reports still get sent. They still get opened, mostly. But there is a version of the next 12 months where a client asks the ChatGPT question, gets a screenshot in response, and quietly starts taking a second meeting with an agency that has a chart.
I do not know how fast that happens. Nobody does. The people giving you a confident timeline on this are guessing with more conviction than I have, and some of them are selling something. What I am fairly sure of is the direction, and that the cost of being early here is a few months of measuring something that turns out to matter less than expected. The cost of being late is a client conversation you cannot win.
The agencies that add this line item in 2026 are the ones renewing in 2027. That is not a prediction about technology. It is a prediction about what happens when one vendor can answer a question and the other one cannot.
Being found was the old job. Getting chosen is the new one.