Method
How the measurement works
The method is deliberately simple, and we would rather explain its limits than oversell it. Here is exactly what we do.
What we ask
We ask the questions your buyers ask, in full sentences. Not keywords. A traveller does not type “best safari Kenya” into an assistant. They ask which operator to trust for a family trip in August, or whether a camp is suitable for a first safari. We write 60 of those questions for your market, with you, before anything is measured.
How many times
Every question is asked three times. AI answers are not consistent: the same model, given the same question a minute later, will often name a different set of businesses. Asking once produces a number that looks precise and is not. Three runs give a rate rather than an anecdote.
Which models, and why the panel differs by market
We use five models per audit. The panel is chosen for the market being measured, because your customers do not all use the same assistant. A German family and a Chinese student are not asking the same software. For mainland China we include a model available there, since ChatGPT is not. Where a market is served mainly by Western assistants, the panel reflects that instead. We tell you which five were used, every time.
The five assistants we ask on safari and educational travel: ChatGPT, Claude, Perplexity, Gemini and Mistral. The first four search the web while they answer; Mistral answers from what it already knows. That contrast is deliberate — it separates what AI can find about you from what it remembers, and the gap between those two is usually the most useful number in the report.
On language schools the fifth model is different: ChatGPT, Claude, Perplexity, Gemini and Qwen. The first four search the web while they answer; Qwen answers from what it already knows, and is included here rather than a European model because a large share of language-school enquiries begin in Chinese.
The eight crawlers we test
For each one we read what your robots.txt says and then make a real request identifying as that crawler, so we can tell a stated policy from an actual one.
OAI-SearchBot (ChatGPT search results) · PerplexityBot (Perplexity answers and citations) · ClaudeBot (Claude web search) · Bingbot (Microsoft Copilot) · Google-Extended (Gemini grounding) · GPTBot (OpenAI model training) · ChatGPT-User (live fetches when someone shares your link) · Applebot-Extended (Apple Intelligence)
Memory and retrieval, reported separately
Some models answer from training data: what they absorbed about you before the conversation started. Others fetch live pages while answering. These produce different results for the same business, and they have opposite fixes.
A high memory score with a low retrieval score means you are well known and unreachable, which is usually a technical problem and often quick to solve. The reverse means your site is readable but little has been written about you, which takes far longer. We report both numbers, and the gap between them is the diagnosis.
The technical audit
Separately from the questions, we probe your site the way a crawler does. We request your pages while identifying as each named AI bot, and record what the server actually returns, because a firewall will often refuse a bot while serving a browser normally. We then check your structured data, whether the page content exists without JavaScript, your sitemaps, your robots.txt, and whether you publish an llms.txt.
Each finding is written up with the effort to fix it and the likely impact, so you can decide what is worth doing.
Limitations
Our figures come from model APIs, not the consumer apps. The apps add memory of previous conversations, personalisation, account context and their own routing, so a real user may see an answer we did not. We cannot measure inside somebody else’s app, and neither can anyone else.
Treat the results as directional and comparative: useful for tracking your own position over time and against named competitors, not as an absolute share of every AI answer given about your industry. The technical findings are different in kind. Those are observed facts about your server, and they are either true or they are not.