webstoree
NOTE

What makes ChatGPT cite a website?

What makes ChatGPT cite a website? A two-million-citation study says it is your buyers' own words, not FAQ blocks or schema. What to change, with a script.

Answer enginesGenerative enginesSearch

A hand pulling one drawer out of a wall of wooden library card catalogue drawers, index cards visible inside

For six months, the same buyer questions were typed into ChatGPT, Claude, Gemini and Google's AI Overviews on a rotating schedule, until there were about two million citations to count. On 28 September 2026 the results went up on arXiv. So what makes ChatGPT cite a website? Mostly this: the page uses the same words your buyers use when they ask. FAQ blocks, schema and page speed, once you account for how big the brand is, barely registered.

The page that gets quoted is the one that sounds like the question, not the one with the longest checklist.

In short:

  • Across about two million citations, the strongest page-level signal was how closely a page's wording matched the questions buyers type.
  • FAQ blocks, schema and page speed looked helpful in pooled data, then vanished once each brand was compared only with itself.
  • Start with your customers' own words: put them in headings and opening lines, high on the page.

What did the two-million-citation study actually measure?

The paper is What Drives Citations in Production Large Language Models? by Ben Moore and Liam Dunne of Discovered Labs, an answer engine benchmarking company. It is worth knowing who wrote it, and also worth knowing what they did, because it is more careful than most of what gets quoted in this field.

The arXiv abstract page for What Drives Citations in Production Large Language Models? by Ben Moore and Liam Dunne, submitted 28 September 2026, with the abstract describing about 2 million citations from ChatGPT, Claude, Google AI and Gemini and the prompt-content alignment result of beta plus 0.37

Screenshot: Moore and Dunne, What Drives Citations in Production Large Language Models? (arXiv 2609.35077), courtesy of arXiv.

They collected roughly 2.1 million citation records from four engines (ChatGPT, Claude, Google AI Overviews and Gemini) across nineteen B2B software companies. Perplexity was dropped for having too few samples. They then crawled 10,042 of the cited pages and measured more than sixty features on each: word count, outbound links, FAQ and summary blocks, author bios, schema, real-user Core Web Vitals, page age, page type, and how closely the page's wording matched the prompts buyers type.

The important design choice is the one most industry studies skip. They compared pages within the same domain. A big, famous brand gets cited a lot and also tends to have FAQ sections, schema and slow pages. Pool everything together and the FAQ section looks like it is doing the work, when it is the brand. Hold the domain fixed and you see what a page-level change is worth on its own.

Then they ran nine statistical methods and would only call an effect real if it passed a five-part test, including a hold-out month the model had not seen. Only a handful of features made it.

Forest chart of standardised coefficients from the study. Prompt-content alignment is far the strongest at plus 0.369. Best paragraph, title and intro similarity and page age are small positives, outbound link count is negative at minus 0.180, and author bio, real-user CLS, FAQ section, page length and answer in top three paragraphs are not significant.

So what makes ChatGPT cite a website, according to the data?

One feature towers over the rest: prompt-content alignment, measured as the overlap between the words and two-word phrases on the page and the full set of prompts buyers were asking. Its standardised coefficient was +0.37, which the authors translate as roughly 30% more citations for a page one standard deviation better aligned, at the typical page. Discovered Labs' own write-up leads with the same number.

The Discovered Labs research page headed What actually drives AI citations: a statistical analysis of 2M AI citations across 10K pages, by Liam Dunne and Ben Moore, with a panel showing prompt-content alignment at plus 0.37 against plus 0.07 for the best on-page signal, FAQ

Screenshot: What actually drives AI citations: a statistical analysis of 2M AI citations across 10K pages, courtesy of Discovered Labs, captured 4 Oct 2026.

A few smaller signals also held up:

  • The best paragraph matters. How closely the single best-matching paragraph, the title and the intro resembled the prompt each had a small positive effect.
  • The answer sits high. The best-matching paragraph had a median position 36% of the way down the page. Engines cite from the top third.
  • Pricing pages carried a large positive effect of their own, even after everything else was controlled. Buyers ask about price, and engines go and find it.
  • Outbound links were negatively associated with citation. That is an association within these nineteen companies, not an instruction to strip out your references.

The authors are honest about the limit: this is observational. It shows that aligned pages are cited more, not that rewriting a page will cause the same lift. But by its authors' account it is the first large, confound-controlled study of production engines, and it points in a clear direction.

Do FAQ blocks, schema and page speed help AI citations?

Not in a way this study could detect, and this is the finding that will annoy people selling them. In pooled data, FAQ sections and schema look helpful and fast pages even look worse. Within the same domain, those effects shrink to nothing or flip. The authors call it Simpson's paradox: large incumbents collect more citations, have slower than median pages and differ from challengers on most page-level signals, so the brand was doing the work all along.

In the final table, FAQ section, author bio, real-user layout shift and page length all fail to reach significance. So does "answer in the top three paragraphs", which is a useful corrective to the instinct to cram everything into the opening.

Domain strength was the other giant. In the authors' tree model, the domain's citation track record carried about six times the importance of the strongest page feature other than alignment. That is the part you cannot fix in an afternoon, though the authors make a fair point: a domain's standing is the accumulated result of years of pages people found useful. It is a stock, built one aligned page at a time.

None of this means schema is pointless. Google still uses structured data to understand pages, and we cover the groundwork in how to show up in Google AI Overviews. It means schema is not the reason one page gets quoted over another.

Will AI Recommend Your Business? Websites Built for AI Search

Why do the famous GEO tricks not transfer?

Because most of them were measured at the wrong end of the pipeline. In July, Olivier Martinez published a critical survey of 45 GEO studies from November 2023 to July 2026. It treats an AI citation as the last step of a long chain: the engine decides to search, crawls and indexes, retrieves candidates, reranks them into a context, and only then cites, quotes and absorbs.

Diagram of nine stages a page passes through before an AI engine cites it: search activation, crawling and indexing, retrieval, reranking and context, then citation, prominence, absorption, fidelity and user behaviour. The first four are highlighted as where relevance and position decide things; citation to absorption is marked as where the famous GEO gains were measured.

The much-quoted "up to 40%" figure comes from the original GEO paper (Aggarwal and colleagues, KDD 2024). The survey points out how it was measured: the top five Google results for a query were handed to the model, one of them was rewritten (adding quotations, statistics or citations), and its share of the answer was compared. A real effect, but only for a page already in the room. It says nothing about getting in.

The survey's conclusions are sober. Topical relevance and position in the context are the most reproducible levers. Generic heuristics transfer poorly between engines. Rewrites aimed at being cited can hurt retrieval, the stage before. And across the whole corpus, no technique showed a stable, long-term, cross-platform effect on being discovered. Two independent pieces of work, one survey and one field study, landing on the same word: relevance.

How can you check whether a page uses your buyers' words?

You can approximate the study's main measure in a few dozen lines. The script below takes a URL and a text file of buyer questions, one per line. It reports a page-level Jaccard score in the spirit of the paper's, and, for each question, the share of its terms that the best single paragraph contains.

// align.mjs: how closely does a page use the words your buyers type?
// Usage: node align.mjs <url> prompts.txt   (one buyer prompt per line)
// Node 18+, no dependencies.
const [url, promptFile] = process.argv.slice(2);
const fs = await import("node:fs");

const STOP = new Set("a an the and or of to in on for is are be it this that with as at by from your you we our can do does how what why which".split(" "));
const words = (s) => s.toLowerCase().match(/[a-z0-9]+/g)?.filter((w) => !STOP.has(w)) ?? [];
// unigrams plus bigrams, as in the paper's lexical Jaccard
const grams = (s) => {
  const w = words(s);
  return new Set([...w, ...w.slice(1).map((x, i) => `${w[i]} ${x}`)]);
};
const jaccard = (a, b) => {
  let inter = 0;
  for (const g of a) if (b.has(g)) inter++;
  return inter / (a.size + b.size - inter || 1);
};
// share of a prompt's terms that one paragraph contains (paragraphs are long, prompts short)
const coverage = (prompt, para) => {
  let hit = 0;
  for (const g of prompt) if (para.has(g)) hit++;
  return hit / (prompt.size || 1);
};

const html = await (await fetch(url)).text();
const main = (html.match(/<main[\s\S]*?<\/main>/i) ?? [html])[0];
const paras = [...main.matchAll(/<(p|li)(?:\s[^>]*)?>([\s\S]*?)<\/\1>/gi)]
  .map((m) => m[2].replace(/<[^>]+>/g, " ").replace(/&[a-z#0-9]+;/gi, " ").trim())
  .filter((t) => t.split(/\s+/).length >= 8);

const prompts = fs.readFileSync(promptFile, "utf8").split(/\r?\n/).filter(Boolean);
const promptGrams = grams(prompts.join(" "));
const page = grams(paras.join(" "));

// for each prompt, which paragraph answers it best, and how far down the page is it?
const rows = prompts.map((p) => {
  const g = grams(p);
  let best = { i: -1, s: 0 };
  paras.forEach((t, i) => {
    const s = coverage(g, grams(t));
    if (s > best.s) best = { i, s };
  });
  return {
    prompt: p,
    score: +best.s.toFixed(3),
    depth: best.i < 0 ? null : +(best.i / Math.max(paras.length - 1, 1)).toFixed(2),
    starts: paras[best.i]?.slice(0, 40),
  };
});

const outbound = [...main.matchAll(/href="https?:\/\/([^/"]+)/g)].filter((m) => !url.includes(m[1])).length;
console.log({ url, paragraphs: paras.length, pageAlignment: +jaccard(page, promptGrams).toFixed(3), outboundLinks: outbound });
console.table(rows);

Be clear about what it is. It is a lexical check, close to the paper's headline measure, not the authors' released pipeline, and it knows nothing about meaning. A paragraph that says "nosnippet" answers "can I stop my site appearing" perfectly well while sharing almost none of its words. That is exactly the point, though: the engine has to retrieve the paragraph before it can understand it.

We ran it on our own post about AI Overviews with six questions a buyer might plausibly type. It was humbling.

Bar chart of the alignment script run on the webstoree post about AI Overviews. For six buyer questions, the best single paragraph contains between 20% and 44% of the question's terms, except "is there special schema for ai overviews" at 78%.

Only one question found a paragraph using most of its words. The answers are all on the page, but phrased the way we talk ("eligible to appear with a snippet") rather than the way buyers ask ("how do I get my website into AI Overviews"). That is a fix measured in sentences, not a rebuild.

What should you change on your pages first?

If what makes ChatGPT cite a website is mostly relevance, the order of work follows. Start with the evidence, not the folklore. This is how the popular levers come out when you put the two papers side by side:

LeverWhat the evidence showsWhat to do
Buyer-language alignmentStrongest page feature, survives every checkRewrite headings and opening lines in the words customers use
Answer high on the pageBest-matching paragraph sits around a third of the way downPut the direct answer in the first screen of each section
Pricing contentLarge positive effect on its ownPublish real prices or ranges, with what drives them
FAQ blocks, author biosNot significant once the domain is held fixedKeep them for readers, not for citations
Schema, Core Web VitalsPooled gains vanish within a domainGroundwork for search; do not expect citation lifts
"Add statistics and quotes" rewritesHelps only pages already in the contextAdd real, sourced numbers because they inform, not as a trick
Being crawlable by the engineA precondition, not a ranking factorAllow OAI-SearchBot if you want to appear in ChatGPT search

That last row deserves one sentence. OpenAI's crawler documentation says sites that block OAI-SearchBot "will not be shown in ChatGPT search answers", and that blocking GPTBot is about training, not search. Plenty of sites block the wrong one.

And where do the buyer's words come from? The authors suggest first-party transcripts: sales calls, support emails, the questions people ask on the phone. Your inbox already contains your best keyword research. If the jargon in this post (AEO, GEO, answer engine) is new, our plain-English glossary has one paragraph on each.

Where this studio stands

We think the study is right about the direction and modest about the size, which is the correct way round. It covers nineteen B2B software companies, not cafes or clinics, and it is observational. But it agrees with a 45-paper survey and with what Google says about AI features: there is no special trick, only relevance, retrieval and a page worth quoting.

So our answer engine optimisation work starts with the questions your customers ask, in their wording, and makes sure a page answers each one high up and in plain sentences. We do not sell FAQ schema as a citation lever, because the best data we have says it is not one. For why an engine might be naming a competitor instead of you, see why AI assistants do not mention your business.

A test for tonight: write down the last five questions a customer asked you, word for word. Then open your website and look for those words. If you cannot find them, neither can ChatGPT.

What this rests on

QUOTE

what will it cost, and when will i hear back?

Get a fixed-price quote.

Tell us what the site has to do. You get a fixed price and a timeline back by email, from the founder who would build it.

A new site, a rebuild, a shop, bookings, an assistant. A sentence or two is plenty.

Reply by email within one business day.