Why an AI quotes some pages and ignores others
Between two pages holding the same knowledge, it is not quality that decides which gets quoted. What decides is whether a sentence can be lifted out of one of them that stands on its own and can be substantiated.
In short
- Sections get quoted, not pages — so the arrangement decides more than the overall quality.
- Six properties are shared by cited sources: an extractable answer, a verifiable detail, visible accountability, currency, technical readability, and thematic surroundings.
- Three reasons rule strong pages out: content only appears through JavaScript, the page is blocked, or the statements are scattered through the prose.
- The verifiable original detail is the strongest single factor — it is the reason a model names you rather than a general source.
When a system builds an answer, it does not choose sources by sympathy or authority but by usability. The question is always the same: does this page supply a passage that answers the question asked, makes sense on its own, and can be substantiated?
The six properties
1. An extractable answer
A section answering the question completely in two to four sentences, without referring to the surrounding text. By far the most important property — and the only one entirely within your control.
2. A verifiable original detail
A figure, a method, a timeframe, an experience with its scope. That is the reason a model names you rather than a general source: you say something that does not exist elsewhere.
3. Visible accountability
Author with role, company with an address, date of last review — visible on the page and in the structured data, with the same value in both. Particularly important on topics touching money, health or law.
4. Traceable currency
A date that is true, and content matching that date. A revision date of yesterday on a text describing a long-superseded position is worse than no date at all.
5. Technical readability
The content sits in the delivered HTML, not only after JavaScript has run. A clean heading hierarchy, and no block on runtime fetches in robots.txt.
6. Thematically related surroundings
A page in a field the website visibly occupies gets treated as an expert source more readily than a single article among twenty unrelated topics. That is why topic clusters work.
Worth knowing
Property two is the actual lever for small companies. On general knowledge you compete with large portals and lose — there is no reason to name you in particular.
With a verifiable original detail you are the source. "In practice, cleaning up a grown contact list takes one to three days" is not common knowledge — it comes from someone who has done it, and that is exactly what makes it quotable. One sentence carrying your own experience weighs more than three paragraphs of summary.
Three reasons strong pages drop out
| Reason | How you spot it | Effort to fix |
|---|---|---|
| Content only appears in the browser | the page source does not contain the text | high – an architecture question |
| Runtime fetches blocked | robots.txt blocks ChatGPT-User and similar | minutes |
| Statements scattered through the prose | no paragraph answers a question alone | medium – rework the paragraphs |
The second reason is the most annoying, because it is fixed in minutes and still occurs frequently — usually as a side effect of a blanket block on "all AI bots".
A pattern that shows up when evaluating cited sources: they are rarely the extensive overview pieces. They are short pages answering exactly one question completely — often ones treated in-house as a by-product.
A 900-word piece with one clear figure and a named scope gets quoted more often than a 2,500-word guide on the same topic where that same figure sits in paragraph eleven. For content planning that means: more sharply bounded individual questions, fewer all-encompassing guides.
What you can concretely do
- Run the extraction test. Copy any paragraph into an empty document: does it alone answer a recognisable question? On most pieces two out of ten paragraphs pass.
- Add original details. For every piece, at least one figure, method or experience with its scope — from your own business, not from a source.
- Put a summary at the top. Four complete statements. Fifteen minutes per piece, with no change to the content.
- Check the runtime fetches. Look in robots.txt whether ChatGPT-User, Perplexity-User and Claude-User are allowed.
- Read the server logs. Are your pages being fetched at all? No access means no citation — and that is a different problem from quotability.
Assess why this page would or would not be quoted by answer engines. Be strict; praise nothing; do not rewrite the text. Page text: [paste text] Further details: - Is the content in the delivered HTML? [yes / no / unknown] - Are author, role and review date visible? [yes / no] - What other field does our website occupy? [topics] Tasks: 1. Check the six properties individually – extractable answer, verifiable original detail, visible accountability, currency, technical readability, thematically related surroundings. For each, give met/partly/not met with a justification from the text. 2. Name the three paragraphs most likely to be quoted, and for each say which question it answers. 3. Name every statement in the text that could equally appear on any other page – that is, containing nothing of your own. 4. Phrase three questions to me whose answers would produce verifiable original details the text currently lacks. Do not invent figures or sources.
In closing
What gets quoted is what can be lifted out and says something that does not exist elsewhere. Both are matters of arrangement and substance, not of a piece's overall quality.
For small companies the lever sits clearly at property two: on general knowledge there is no reason to name you. With your own figure and a named scope, you are the source.
Common questions
Why does an AI quote some pages and not others?
Because it selects sections, not pages. What gets quoted is what can be lifted out and stays understandable on its own, contains a verifiable detail, makes clear who stands behind it, is current and technically readable — and sits in a field the website visibly occupies.
What is the strongest single factor?
A verifiable original detail: a figure, a method, an experience with a named scope. On general knowledge there is no reason to name a small company in particular — with an original detail it is the source.
Why does a substantively strong page sometimes never get quoted?
For three reasons: the content only appears through JavaScript and is not in the delivered HTML; the runtime fetches are blocked in robots.txt, often as a side effect of a blanket block on all AI bots; or the statements are spread through the prose so no paragraph answers a question alone.
Do long or short pieces get quoted more often?
In practice, sharply bounded short ones. A piece of around 900 words with a clear figure and a named scope gets quoted more often than an extensive guide where the same figure sits in paragraph eleven — because the short piece is easier to lift out.
How do you check your own quotability?
With the extraction test: copy any paragraph into an empty document and read it. Does it alone answer a recognisable question completely? On most pieces around two out of ten paragraphs pass — that share is the most informative metric for quotability.
Marketing that sets itself up
The Studio Engine beta is live. Claim your spot and help shape it from the start.
Join the beta →