Measuring visitors from AI answers: what the analytics give you
Visitors from answer engines are few and distinctive: they stay longer, bounce less and come more directly to the point. Capturing them cleanly is possible — but only in part, and that belongs in the report.
In short
- AI traffic is recognised by the referrer — the domains of the chat and answer services appear as their own referral sources.
- Part of it is fundamentally uncapturable: someone who reads the answer and then types your company name appears as direct traffic.
- The server log is the second and often more informative source: it shows whether your pages are being read by AI crawlers at all.
- Three numbers belong in the report — visitors per source, crawler requests per page, and time on page compared with the rest.
The numbers are small. In most small companies the share sits in the low single-digit percentages. It is interesting nonetheless, because it looks different from the rest of the traffic — and because it is growing.
How to recognise it
Visitors arriving from a generated answer generally carry the respective service's domain as the referrer. In analytics they therefore appear as their own source — often under "referral" rather than "search".
| Appears as | Origin | Note |
|---|---|---|
| chatgpt.com | link from a ChatGPT answer | usually unambiguous |
| perplexity.ai | source click in Perplexity | highest volume of the three |
| gemini.google.com | link from Gemini | small, partly unattributed |
| claude.ai | link from a Claude answer | usually unambiguous |
| copilot.microsoft.com | link from Copilot | partly as a Bing referral |
| (direct) | answer read, name typed afterwards | not attributable |
Setting it up in four steps
- Create a channel group. In your analytics, define a group called "AI answers" collecting the referrers listed above. Otherwise they hide among twenty other referrals.
- Do not exclude the referrers. Some analytics tools filter unknown referrals automatically. Check whether these domains are affected.
- Record the landing page. The most important additional detail: which page was linked as a source? From that follows which content is actually citable.
- Read the server logs. Filter on the crawler identifiers: GPTBot, OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, ClaudeBot, Claude-User, Google-Extended, Bingbot.
Worth knowing
At small volumes the server log is the more informative source. It does not show how many people came — but it shows whether your pages are being read at all, which ones, and how often.
Particularly useful is the distinction by identifier: requests from identifiers with "User" in the name arise because someone has just asked a question your page might answer. They are therefore a more direct signal than the advance collectors — even when not a single click follows.
Why these visitors look different
Three patterns show up nearly everywhere there is enough data:
Considerably longer time on page. Someone arriving from an answer has already been through a selection and reads on deliberately.
Fewer pages per visit. They come for one particular statement, not to browse. Few pages is not a bad sign here.
More often landing directly on deep pages. Not the home page but an individual piece, or a section of it. That is exactly why every deep page has to make sense on its own — it is frequently the first contact.
The most useful analysis is not the total but the list of landing pages. It shows which of your content actually serves as evidence — and it surprises people regularly.
Often it is not the elaborate overview pieces but short, clearly answered individual questions. A section answering one concrete question completely in four sentences gets cited more often than a 2,000-word guide on the same topic. That list is therefore a better content plan than any topic collection.
Where the measurement ends
Three things cannot be measured, and that belongs in every report:
- Mentions without a click. Being named in an answer and not clicked leaves no trace. That is probably the larger part of the effect.
- Mentions in systems that do not link. Some answers name names in the running text without linking.
- The delay. Someone reads an answer in March and enquires in June. That chain is unreconstructable.
The only proxy against this is the question in the first conversation: "How did you come across us?" Imprecise, but it captures exactly the cases no measurement reaches.
Help me evaluate traffic from AI answers properly. The data: - Analytics tool: [name] - Period: [details] - Visitors from AI referrers (per source and landing page): [paste table] - Time on page and pages per visit, overall and for this group: [values] - Crawler requests from the server log (identifier, page, count): [paste table] Tasks: 1. Name the landing pages linked most often as a source, and what they have in common – as far as the data shows. 2. Compare this group's time on page and pages per visit with the overall average. Interpret the result. 3. Name pages fetched frequently by crawlers but producing no referral traffic – and what that can mean. 4. Phrase three sentences for our monthly report that honestly say what these numbers show and what they do not. Do not invent values. Where data is missing, say which.
In closing
The measurement is incomplete and still worth doing — above all because the list of cited landing pages shows directly which content is citable.
Three numbers are enough in the report: visitors per AI source, crawler requests per page, time on page compared. Plus a fourth sentence noting that a substantial part of the effect shows up as direct traffic and is therefore not in those numbers.
Common questions
How do you measure visitors coming via AI answers?
Through the referrer in your analytics: the chat and answer services' domains appear as their own referral sources and can be grouped into an "AI answers" channel. The server log additionally shows which AI crawlers fetched which pages — at small volumes, often the more informative source.
Why is part of AI traffic unmeasurable?
Because many people read the answer, remember the name and visit the site later directly. Such visits appear as direct traffic and cannot be attributed. The same applies to mentions without a link and to enquiries that follow months later.
How does this traffic differ from the rest?
Through three patterns: considerably longer time on page, fewer pages per visit, and more frequent entry directly on a deep page rather than the home page. Few pages per visit is not a bad sign here — these visitors come for one particular statement, not to browse.
Which crawler identifiers should you track?
GPTBot, OAI-SearchBot and ChatGPT-User, PerplexityBot and Perplexity-User, ClaudeBot and Claude-User, Google-Extended and Bingbot. The distinction is useful: identifiers with "User" in the name fetch a page because someone has just asked a matching question — a more direct signal than the advance collectors.
What do you do with the analysis?
The list of landing pages linked as sources is the most useful part: it shows which content is citable. Often those are not the extensive guides but short, clearly answered individual questions — which should feed directly into next quarter's content planning.
Marketing that sets itself up
The Studio Engine beta is live. Claim your spot and help shape it from the start.
Join the beta →