Perplexity, ChatGPT, Gemini: how the answer engines really search

Anyone optimising for answer engines should know that there is no such thing as “the” answer engine. The systems differ on the point that decides visibility: where they take their sources from and how many of them they name.

One source of light passes through three different prisms, each casting a differently shaped pattern

In short

  • All three systems fall back on web search for current questions — but on different indexes and with different numbers of sources.
  • Perplexity names the most sources and is the most accessible for small providers. Gemini names the fewest and leans heavily on Google's results.
  • Two things work everywhere: a clear, extractable answer per section, and a page readable without JavaScript.
  • Success can only be measured repeatably, not precisely — a fixed list of questions, asked quarterly across all systems.
Note These systems change quickly and do not publish their selection procedures. The assessments below rest on observable behaviour and providers' public statements, not on disclosed mechanisms. They are a snapshot from August 2026.

The shared principle

For current questions, all three work to the same pattern: they break the question into search queries, fetch results, read some of the pages found, and build an answer with source references from them.

From that follows the most important point for your own visibility: what does not appear in conventional search rarely appears in the answers either. Answer engine optimisation without the groundwork of search engine optimisation does not work.

Where they differ

PerplexityChatGPTGemini
Source referencesmany, prominentfew to moderatefew, restrained
Searches by defaultnearly alwayswhen currency is neededwhen currency is needed
Draws onown index plus partnersown crawler plus partnersheavily on Google results
Small providerscomparatively accessiblemoderateharder, follows the ranking
CrawlersPerplexityBot, Perplexity-UserGPTBot, OAI-SearchBot, ChatGPT-UserGooglebot, Google-Extended
Referral traffichighestmoderatelowest

Worth knowing

The number of sources named directly determines how realistic your own visibility is. A system naming eight sources has eight places to give; one naming two has two.

For small providers a practical order follows: the route into Perplexity answers is shorter than the route into Gemini answers — not because content is judged differently there, but simply because there is more room. If you want to measure whether AEO work is having an effect, that is where you will see something first.

The crawlers and their roles

One point that regularly gets confused: there are two kinds of request, and they are controlled separately.

Training and index crawlers

Collect content in advance — for training or for a search index of their own. Examples: GPTBot, PerplexityBot, ClaudeBot, Google-Extended. They visit your page independently of any specific question.

Runtime fetches

Fetch a page exactly when someone has asked a question it might answer. Examples: ChatGPT-User, Perplexity-User, OAI-SearchBot. They come in small numbers but with a direct connection to a real question.

The distinction matters practically: blocking all AI crawlers prevents not only use in training but also being named as a source in answers. Blocking only the training crawlers keeps you visible in the answers — a compromise many choose deliberately.

Tip Check your server logs for which of these names actually appear. That is the only sound statement about whether your pages are being read at all — and it costs nothing beyond filtering by user agent.

What works on all three

Despite the differences, four measures take effect everywhere.

  1. Answer first, per section. The first two sentences under a heading have to answer the question completely. That is exactly the passage that gets extracted.
  2. Readable without JavaScript. Content that only appears in the browser is not captured by some fetches. That is the most common technical reason for exclusion of all.
  3. Verifiable original details. Figures, methods, named experience. Systems favour pages with specific content over summaries when choosing sources.
  4. Visible accountability. Author with role, date of last review, company with an address. Marked up as structured data, visible on the page.
From practice

A result that shows up regularly in reviews: pages with an "in short" section at the very top get cited noticeably more often than those without — with otherwise identical content.

The reason is mechanical rather than editorial: that section supplies exactly the form a system needs for an answer — short, complete, with no reference to surrounding text. It is the cheapest AEO measure there is: fifteen minutes per piece, with no change to the actual content.

How to measure it

There is no console. What there is, is a repeatable procedure — and repeatable matters more than precise.

  1. Fix twenty questions. From your field, phrased the way a customer would put them. That list stays unchanged for a year.
  2. Ask them quarterly across all three systems. Without a signed-in account, so personalisation plays no part.
  3. Note three things: are you named? which competitors are named? which sources serve as evidence?
  4. Cross-check the server logs. Which crawlers came, which pages did they fetch.

The third line is the most useful. The cited sources show what kind of page the system considers citable — and that is a more concrete instruction than any general recommendation.

Prompt
Help me build a fixed question list for checking visibility in
answer engines.

Our situation:
- What we offer: [offering]
- Audience: [as narrow as possible]
- Region and language: [details]
- Topics we have content on: [list]

Tasks:
1. Phrase 20 questions the way our audience would actually ask
   them – in their language, not our technical vocabulary.
   Distribute them like this: 8 questions describing the problem,
   8 on method, 4 on comparing providers.
2. For each question, mark whether we already have content on it
   (based on my topic list) or not.
3. Name the 5 questions where being mentioned would be worth most
   to us, and justify that.
4. Propose a simple table for recording the results each quarter.

Do not answer the questions yourself, and do not tell me whether
we are currently mentioned – I will check that in the systems.

In closing

The systems differ measurably in how many sources they name and how easily small providers get there. For practical work that changes the order, not the measures: all four effective steps work everywhere.

If time is short, do two things first — put a short summary at the top of every piece, and check that the page is fully readable without JavaScript. After that the question list is worth setting up, so you can see whether anything is moving at all.

Common questions

How do Perplexity, ChatGPT and Gemini differ in searching?

Mainly in the number of sources they name and where the results come from. Perplexity searches nearly always, names many sources prominently, and is the most accessible for small providers. ChatGPT searches when currency is needed and names fewer sources. Gemini leans heavily on Google's results and names the fewest — the route there is longest for small providers.

Do you have to optimise for each system separately?

No. Four measures work on all of them: a complete, extractable answer in the first two sentences of every section, a page readable without JavaScript, verifiable original details rather than summaries, and visible accountability through author, role and review date.

What is the difference between GPTBot and ChatGPT-User?

GPTBot collects content in advance, independently of any specific question. ChatGPT-User fetches a page exactly when a user has asked a question it might answer. Both can be controlled separately in robots.txt — blocking only the advance collectors keeps you visible as a source in answers.

Should you block AI crawlers?

If you sell content or earn from ad impressions, there are good reasons to block at least the advance collectors. If you sell projects, blocking entirely loses you more than it gains — being named as a source reaches someone in the middle of looking for a solution, with a recommendation that did not come from you.

How do you measure visibility in answer engines?

With a fixed list of twenty questions asked quarterly across all systems without a signed-in account. Note three things: whether you are named, which competitors are named, and which sources serve as evidence. Server logs additionally show which AI crawlers actually fetched your pages.

Marketing that sets itself up

The Studio Engine beta is live. Claim your spot and help shape it from the start.

Join the beta →
← Back to overview