Known Is Not the Same as Recommended
A study of 1,608 buyer questions found AI assistants could identify all 40 companies by name, yet never shortlisted 9 of them. The wording of the question decided who got named.

Here is the result that stopped me cold. A UK agency put questions about 40 British B2B software companies to ChatGPT, Gemini, Google AI Overviews and Google AI Mode. Ask about any one of those firms by name, and all four assistants knew exactly who it was. Forty out of forty. Perfect recall.
Then the researchers stopped naming names and started asking the way a buyer actually asks. Best tool for this job. Who should I shortlist. What are the alternatives. Nine of those same forty companies were never mentioned. Not once, across 24 shortlist answers each.
Known. Recognised. Correctly described. And completely absent from the moment that decides the purchase.
The setup, and its limits
The study comes from Recrawled, a UK agency that sells answer-engine optimisation, which is worth saying out loud before anything else. They ran 1,608 buyer-style questions on a single day (4 September 2026), signed out, from a UK location, and published on 29 September. One day. Forty firms. A vendor with an obvious interest in the finding.
So treat the exact percentages as a snapshot rather than a law of physics. The pattern underneath is still the most useful public data I have seen on how AI assistants actually pick who to recommend.
Four ways to ask, four different worlds
This is the part I keep rereading. Same companies, same day, same four assistants. The only variable was how the question was phrased:
- Shortlist question that mentioned the UK: 63% of firms got named
- Same question framed for a mid-sized buyer, no country mentioned: 43%
- Asking for alternatives to the category leader: 13%
- Asking the leader versus a named rival: 0%
Visibility went from roughly two in three, to zero, purely on phrasing. Nothing about the companies changed between those four questions. Nothing about their products, their customers, their websites.
Think of it as a nightclub where the bouncer works from a different guest list depending on how you word your request at the door. Say one thing and you are waved through. Rephrase it slightly and you were never on any list at all.
Your brand does not have a ranking in AI search. It has a hit rate, and that hit rate changes with every way a customer phrases the question.
The review badges did not save anyone
Here is the finding that should rattle a lot of marketing plans. The researchers checked whether a firm's combined G2 and Capterra review count predicted whether an AI assistant would name it. The correlation came back at -0.06.
That is statistical noise. Not a weak positive relationship, not a negative one, just nothing. Years of chasing review volume, badges and star counts bought no reliable seat on the AI shortlist in this sample. (Correlation is not causation and forty firms is a small pool, so do not burn your review programme over one study. But do stop assuming it is the whole answer.)
Two failure modes nobody budgets for
First: 8 of the 40 firms had their own pages cited as a source in answers that then recommended somebody else. You wrote the reference material, the assistant read it, found it credible enough to link, and used it to make the case for a competitor. It is lending your notes to the classmate who then gets picked ahead of you.
Second: about one in three firms got at least one answer that was about the wrong business entirely. Same-name confusion. If your company shares a name with a pub, a band, a clinic or a footballer, that is not a hypothetical, it is a live leak in your pipeline that no analytics dashboard will show you.
There was also a surface split worth noting: YouTube showed up in roughly 25% of AI Overview answers and in zero of 360 ChatGPT answers. The report describes that metric two different ways, so hold it loosely, but the direction is intuitive. Google cites video. ChatGPT, in this test, did not.
And the ground underneath is still moving
On 2 October, Cloudflare shipped a Web Search API inside its AI Gateway. An agent calls one endpoint and picks a provider: Ceramic.ai (its own index, 40 billion plus pages) at $0.25 per 1,000 searches, Linkup at $5, Exa at $7. Providers have to be verified bots that identify themselves and obey robots.txt, and every result must link its source.
Two things follow from that. Grounding an agent in the live web is now priced in fractions of a cent (twenty searches at the default provider costs about half a US cent, before model tokens). And there is no single index sitting behind the phrase "AI answers". There is a 28-fold price range across three different maps of the web, and an agent builder chooses one.
Which means you can be the obvious answer inside one assistant and functionally invisible in another, not because you did anything differently, but because they are reading different maps.
What I would actually do with this
- Measure per question and per platform. A single "AI visibility score" averages away the only thing that matters. 63% and 0% do not average into anything useful.
- Test the four shapes. Category question, buyer-situation question, alternatives-to-the-leader question, head-to-head question. Those behave like four separate channels.
- Run the wrong-business check. Ask each assistant about your brand name cold and see who it thinks you are.
- Check that crawlers can reach you. If the indexes feeding these answers cannot fetch your pages, nothing else on this list matters.
For twenty years, being known was the goal. Get the brand recognised, get the pages indexed, and recommendation followed. The assistant layer has quietly split those two things apart. Recognition is now table stakes, and all forty firms in this study had it. Recommendation is a separate game, played per question, per platform, and apparently won by something other than review count.
Most companies are still measuring the first thing. The buying decision has moved to the second.
One signal a day. No noise.
A 3-minute read when something genuinely shifts in AI, automation, or defense tech. Free, most weekdays.
Free, most weekdays. No spam, unsubscribe anytime.Sources
- PPC Land - ChatGPT, Gemini and Google never name 9 of 40 UK software firms in shortlists - https://ppc.land/chatgpt-gemini-and-google-never-name-9-of-40-uk-software-firms-in-shortlists/
- PPC Land - Cloudflare Web Search for AI agents costs $0.25 to $7 per 1,000 requests - https://ppc.land/cloudflare-web-search-for-ai-agents-costs-0-25-to-7-per-1-000-requests/
Quick answers
What did the AI shortlist study actually test?
Recrawled, a UK answer-engine optimisation agency, put 1,608 buyer-style questions about 40 UK B2B software firms to ChatGPT, Gemini, Google AI Overviews and Google AI Mode on 4 September 2026, signed out and from a UK location. The report was published on 29 September 2026.
Why does the wording of a question change which companies get recommended?
The study did not establish a cause, but it measured the effect clearly. Firms were named in 63% of shortlist questions that mentioned the UK, 43% of mid-sized buyer questions with no country, 13% of alternatives-to-the-leader questions and 0% of head-to-head comparisons between a leader and a named rival. Same firms, same day, different phrasing.
Do software review sites help a brand get named by AI assistants?
In this sample, no. The correlation between a firm's combined G2 and Capterra review count and whether AI assistants named it was -0.06, which is effectively zero. That is one study of 40 firms on one day, so it is a warning against relying on review volume alone, not proof that reviews are worthless.
Can my site be cited by an AI answer that recommends a competitor?
Yes. 8 of the 40 firms in the study had their own pages cited as sources in answers that went on to recommend a different company. Separately, about one in three firms got at least one answer about a completely different business with a similar name.