How to Get Your Site Cited by ChatGPT and Perplexity (GEO for Small Sites)

Here's the statistic that should reorganise your priorities: research on ChatGPT's citation behaviour found that around 90% of the pages it cites are not in Google's top 20 for the equivalent query. Read that again from a small site's perspective. The channel where domain authority decides everything is being joined by one where it barely registers — where a specialist page from an unknown domain gets cited next to (or instead of) the incumbents.

That's what makes generative engine optimisation (GEO) interesting for sites that can't win link wars. But most GEO advice is abstract theory written for enterprise brands. This is the tactical version for small sites: what AI engines actually select for, and what to change on your pages this week.

TL;DR

  • AI assistants don't rank you — they retrieve, then quote. Your job is to be the easiest source to quote accurately.
  • The highest-leverage change is fact density: the original GEO research found that adding statistics, quotations, and cited sources boosts a page's AI visibility by 30–40% — with the biggest gains for lower-ranked sites.
  • Structure for extraction: answer-first paragraphs, question-matching H2s, tables, tight self-contained passages.
  • Cover the plumbing: don't block GPTBot/PerplexityBot/ClaudeBot in robots.txt, get indexed in Bing (cheap, and it drives Copilot visibility), keep schema clean, ship an llms.txt.
  • Platforms differ: ChatGPT links sources like footnotes; Gemini mentions brands but rarely links. Optimise for being quotable and nameable, not just clickable.
  • Measure it: set up an AI Traffic channel in GA4 before you start, so you can see what's working.

How AI Citation Actually Works (in 60 Seconds)

When someone asks ChatGPT or Perplexity a question with a live answer, the assistant runs web retrieval, querying an index, pulling a handful of candidate pages, and synthesising an answer that cites the passages it used. Perplexity runs its own crawl. What sits behind ChatGPT's retrieval is no longer something anyone outside OpenAI can state with confidence.

â„šī¸
Updated August 2026 This guide used to say ChatGPT's retrieval draws heavily on Bing. That was a fair description in 2023 and 2024, but it is no longer safe to state as fact. OpenAI operates its own crawler layer, and independent tests through 2025 and 2026 have repeatedly found ChatGPT citing pages that appear in Google's index and not Bing's. Nothing below depends on the answer, since passage-level selection works the same way either way, but plan for opacity rather than for a specific dependency. Our post on the Search Console AI report covers why Bing's citation data is still worth tracking in its own right, just not as a proxy for ChatGPT.

Two consequences follow, and all GEO tactics derive from them:

  1. Selection happens at the passage level, not the domain level. The model is looking for the chunk of text that answers the question cleanly. This is why unknown domains get cited: a perfect passage on a DR-15 site beats a rambling one on a DR-90 site.
  2. The model cites what it can lift. Content that's self-contained, factual, and unambiguous survives synthesis. Content that requires three paragraphs of context, or hedges everything, doesn't get quoted — there's nothing to quote.

Worth knowing what's at stake before you invest in any of this. A randomised field experiment found that displaying an AI Overview cut organic clicks by 39.8%, and that the visits lost were statistically indistinguishable from the ones that survived. Citation is not a consolation prize for traffic you were keeping anyway: the clicks being absorbed were representative, and they concentrate on exactly the informational queries these tactics target.

The Playbook

1. Make every key page fact-dense

The original GEO research (Aggarwal et al., the paper that named the field) tested which content changes increase AI visibility. The winners were consistent: adding statistics, adding quotations, and citing sources — improvements of 30–40% in how often and how prominently content gets used, with the strongest effect for pages that don't rank well in classic search. Keyword stuffing, by contrast, did nothing.

In practice: replace vague claims with numbers ("most sites fail INP" → "roughly 4 in 10 sites fail INP's 200 ms threshold"); name your sources inline; include at least one stat-per-section on commercial pages. And the compounding move — publish original numbers. Your own study, survey, benchmark, or client data, however small, makes you the primary source other pages cite, which makes you the source models cite. A small site with one original dataset outranks its authority class in AI answers indefinitely.

2. Structure for extraction, not just reading

  • Answer-first sections. H2 phrased as the question users ask; first sentence under it is the complete answer; elaboration after. (You've seen this pattern across our articles — it targets featured snippets and AI citation with the same stroke.)
  • Self-contained passages. Each section should make sense quoted alone. Avoid "as mentioned above" dependencies in your key answers.
  • Tables and lists for comparisons. Models extract structured data far more reliably than prose, and comparison queries ("X vs Y") are among the most common assistant questions.
  • A TL;DR or key-takeaways block. You're literally pre-writing the synthesis; assistants frequently lift it near-verbatim.

3. Clear the technical path

  • robots.txt audit, today. Check you're not blocking GPTBot, OAI-SearchBot, PerplexityBot, or ClaudeBot — plenty of sites blocked AI crawlers during the 2023 training-data backlash and forgot, which now means invisibility in AI search. Blocking training bots while allowing search/retrieval bots is a legitimate middle position; know which you're doing.
  • Bing is still worth doing. Not because ChatGPT provably depends on it, which is no longer a claim anyone outside OpenAI can make, but because Bing indexing drives Copilot visibility and unlocks the Citation Share data in Bing Webmaster Tools. Verify your site there, submit sitemaps, and check your key pages are actually indexed. Ten minutes of work most sites have never done.
  • Schema markup feeds the knowledge layer that helps engines resolve who you are and what your pages claim — the machine-readability work transfers directly to AI retrieval. On a large site that means consistent entity IDs across templates rather than page-by-page markup, which is the subject of schema markup at scale.
  • llms.txt for the retrieval engines that read it (Perplexity, Claude, agents). Fifteen minutes, small upside, zero risk.

4. Be present where models look when they're unsure

Assistants triangulate entities across the open web. For brand-adjacent queries ("best X tools", "X agency reviews"), they draw disproportionately on third-party surfaces: Reddit threads, industry listicles, review platforms, Wikipedia/Wikidata where applicable. For a small site, a mention in two or three credible "best of" roundups and genuinely helpful (not astroturfed) participation in your niche's Reddit/forum spaces does more for AI-answer presence than another month of link building. This is classic digital PR with a new payoff curve.

5. Know the platform differences

How the major AI assistants cite, and what to optimise for each
Platform Behaviour What to optimise
ChatGPT Cites linked sources footnote-style (~87% of answers with retrieval); brand mentions rarer Quotable passages, Bing indexing, fact density
Perplexity Heaviest citation culture; every claim sourced; own crawler Same as above + llms.txt; fast, clean pages help its crawler
Gemini Mentions brands in most answers but links only ~20% of the time Entity clarity: consistent naming, schema, third-party corroboration — being nameable matters more than clickable
Claude Retrieval on request; reads llms.txt Clean structure, llms.txt

The strategic read: optimise pages to be quotable (ChatGPT/Perplexity) and your brand entity to be nameable (Gemini). They're different assets — passages vs reputation — and small sites usually neglect the second.

6. Measure, then double down

Before any of this, set up the AI Traffic channel in GA4 so you have a baseline. The landing-page report inside that channel is your feedback loop: it shows which pages assistants already cite. Those pages reveal your citable pattern — replicate it. And remember the traffic is only the visible part; brand mentions inside answers (the zero-click layer) don't show in analytics but shape buying decisions. Checking your key queries manually in ChatGPT/Perplexity once a month is unglamorous and effective.

The Quality Caveat

Every tactic above assumes the content underneath is worth citing. Fact density added to derivative content is lipstick — models increasingly cross-check claims, and the same low-value patterns Google demotes make content unquotable too. GEO isn't an alternative to making genuinely useful pages; it's packaging that lets AI engines find the value that's there.

Frequently Asked Questions

How do I get my website cited by ChatGPT?

Make key pages fact-dense (statistics, sources, quotable self-contained passages), structure them answer-first, ensure your site is indexed in Bing and doesn't block OpenAI's crawlers, and build third-party mentions in your niche. Citation selection happens at the passage level, so content quality per section matters more than domain authority.

Does domain authority matter for AI citations?

Far less than in Google. Research shows ~90% of ChatGPT's citations come from pages outside Google's top 20 for the same query. AI engines select the best passage, not the strongest domain — the structural opportunity for small sites.

What is GEO (generative engine optimization)?

The practice of optimising content so AI engines — ChatGPT, Perplexity, Gemini, AI Overviews — retrieve and cite it in their answers. It overlaps with SEO (crawlability, structure, quality) but optimises for passage extraction and entity recognition rather than ranking position.

Should I block AI crawlers from my site?

If you want AI-search visibility, no — blocking GPTBot or PerplexityBot removes you from retrieval. Some sites block training crawlers while allowing search bots; just make sure your robots.txt reflects a decision rather than a 2023 default you forgot about.

Want Your Site Legible to Every Engine — Traditional or Generative?

GEO rewards the same foundation we build for search: machine-readable structure, clean crawl paths, and content with something original in it.

About the Author

Shakur Abdirahman
Technical SEO Specialist
Shakur is a Technical SEO Specialist who implements directly inside WordPress and Shopify environments. He helps small and large sites become legible to every engine — traditional search and generative assistants alike — through machine-readable structure, clean crawl paths, and content built to be cited.

More about the author →