How to Get Your Site Cited by ChatGPT and Perplexity (GEO for Small Sites)

Here's the statistic that should reorganise your priorities: research on ChatGPT's citation behaviour found that around 90% of the pages it cites are not in Google's top 20 for the equivalent query. Read that again from a small site's perspective. The channel where domain authority decides everything is being joined by one where it barely registers — where a specialist page from an unknown domain gets cited next to (or instead of) the incumbents.

That's what makes generative engine optimisation (GEO) interesting for sites that can't win link wars. But most GEO advice is abstract theory written for enterprise brands. This is the tactical version for small sites: what AI engines actually select for, and what to change on your pages this week.

TL;DR

  • AI assistants don't rank you — they retrieve, then quote. Your job is to be the easiest source to quote accurately.
  • The highest-leverage change is fact density: the original GEO research found that adding statistics, quotations, and cited sources boosts a page's AI visibility by 30–40% — with the biggest gains for lower-ranked sites.
  • Structure for extraction: answer-first paragraphs, question-matching H2s, tables, tight self-contained passages.
  • Cover the plumbing: don't block GPTBot/PerplexityBot/ClaudeBot in robots.txt, get indexed in Bing (ChatGPT's retrieval leans on it), keep schema clean, ship an llms.txt.
  • Platforms differ: ChatGPT links sources like footnotes; Gemini mentions brands but rarely links. Optimise for being quotable and nameable, not just clickable.
  • Measure it: set up an AI Traffic channel in GA4 before you start, so you can see what's working.

How AI Citation Actually Works (in 60 Seconds)

When someone asks ChatGPT or Perplexity a question with a live answer, the assistant runs web retrieval — search queries against an index (ChatGPT's retrieval draws heavily on Bing; Perplexity runs its own crawl) — pulls a handful of candidate pages, and synthesises an answer, citing the passages it used.

Two consequences follow, and all GEO tactics derive from them:

  1. Selection happens at the passage level, not the domain level. The model is looking for the chunk of text that answers the question cleanly. This is why unknown domains get cited: a perfect passage on a DR-15 site beats a rambling one on a DR-90 site.
  2. The model cites what it can lift. Content that's self-contained, factual, and unambiguous survives synthesis. Content that requires three paragraphs of context, or hedges everything, doesn't get quoted — there's nothing to quote.

The Playbook

1. Make every key page fact-dense

The original GEO research (Aggarwal et al., the paper that named the field) tested which content changes increase AI visibility. The winners were consistent: adding statistics, adding quotations, and citing sources — improvements of 30–40% in how often and how prominently content gets used, with the strongest effect for pages that don't rank well in classic search. Keyword stuffing, by contrast, did nothing.

In practice: replace vague claims with numbers ("most sites fail INP" → "roughly 4 in 10 sites fail INP's 200 ms threshold"); name your sources inline; include at least one stat-per-section on commercial pages. And the compounding move — publish original numbers. Your own study, survey, benchmark, or client data, however small, makes you the primary source other pages cite, which makes you the source models cite. A small site with one original dataset outranks its authority class in AI answers indefinitely.

2. Structure for extraction, not just reading

  • Answer-first sections. H2 phrased as the question users ask; first sentence under it is the complete answer; elaboration after. (You've seen this pattern across our articles — it targets featured snippets and AI citation with the same stroke.)
  • Self-contained passages. Each section should make sense quoted alone. Avoid "as mentioned above" dependencies in your key answers.
  • Tables and lists for comparisons. Models extract structured data far more reliably than prose, and comparison queries ("X vs Y") are among the most common assistant questions.
  • A TL;DR or key-takeaways block. You're literally pre-writing the synthesis; assistants frequently lift it near-verbatim.

3. Clear the technical path

  • robots.txt audit, today. Check you're not blocking GPTBot, OAI-SearchBot, PerplexityBot, or ClaudeBot — plenty of sites blocked AI crawlers during the 2023 training-data backlash and forgot, which now means invisibility in AI search. Blocking training bots while allowing search/retrieval bots is a legitimate middle position; know which you're doing.
  • Bing is not optional. ChatGPT's web retrieval leans on Bing's index. Verify your site in Bing Webmaster Tools, submit sitemaps, and check your key pages are actually indexed. Ten minutes of work most sites have never done.
  • Schema markup feeds the knowledge layer that helps engines resolve who you are and what your pages claim — the machine-readability work transfers directly to AI retrieval.
  • llms.txt for the retrieval engines that read it (Perplexity, Claude, agents). Fifteen minutes, small upside, zero risk.

4. Be present where models look when they're unsure

Assistants triangulate entities across the open web. For brand-adjacent queries ("best X tools", "X agency reviews"), they draw disproportionately on third-party surfaces: Reddit threads, industry listicles, review platforms, Wikipedia/Wikidata where applicable. For a small site, a mention in two or three credible "best of" roundups and genuinely helpful (not astroturfed) participation in your niche's Reddit/forum spaces does more for AI-answer presence than another month of link building. This is classic digital PR with a new payoff curve.

5. Know the platform differences

How the major AI assistants cite, and what to optimise for each
Platform Behaviour What to optimise
ChatGPT Cites linked sources footnote-style (~87% of answers with retrieval); brand mentions rarer Quotable passages, Bing indexing, fact density
Perplexity Heaviest citation culture; every claim sourced; own crawler Same as above + llms.txt; fast, clean pages help its crawler
Gemini Mentions brands in most answers but links only ~20% of the time Entity clarity: consistent naming, schema, third-party corroboration — being nameable matters more than clickable
Claude Retrieval on request; reads llms.txt Clean structure, llms.txt

The strategic read: optimise pages to be quotable (ChatGPT/Perplexity) and your brand entity to be nameable (Gemini). They're different assets — passages vs reputation — and small sites usually neglect the second.

6. Measure, then double down

Before any of this, set up the AI Traffic channel in GA4 so you have a baseline. The landing-page report inside that channel is your feedback loop: it shows which pages assistants already cite. Those pages reveal your citable pattern — replicate it. And remember the traffic is only the visible part; brand mentions inside answers (the zero-click layer) don't show in analytics but shape buying decisions. Checking your key queries manually in ChatGPT/Perplexity once a month is unglamorous and effective.

The Quality Caveat

Every tactic above assumes the content underneath is worth citing. Fact density added to derivative content is lipstick — models increasingly cross-check claims, and the same low-value patterns Google demotes make content unquotable too. GEO isn't an alternative to making genuinely useful pages; it's packaging that lets AI engines find the value that's there.

Frequently Asked Questions

How do I get my website cited by ChatGPT?

Make key pages fact-dense (statistics, sources, quotable self-contained passages), structure them answer-first, ensure your site is indexed in Bing and doesn't block OpenAI's crawlers, and build third-party mentions in your niche. Citation selection happens at the passage level, so content quality per section matters more than domain authority.

Does domain authority matter for AI citations?

Far less than in Google. Research shows ~90% of ChatGPT's citations come from pages outside Google's top 20 for the same query. AI engines select the best passage, not the strongest domain — the structural opportunity for small sites.

What is GEO (generative engine optimization)?

The practice of optimising content so AI engines — ChatGPT, Perplexity, Gemini, AI Overviews — retrieve and cite it in their answers. It overlaps with SEO (crawlability, structure, quality) but optimises for passage extraction and entity recognition rather than ranking position.

Should I block AI crawlers from my site?

If you want AI-search visibility, no — blocking GPTBot or PerplexityBot removes you from retrieval. Some sites block training crawlers while allowing search bots; just make sure your robots.txt reflects a decision rather than a 2023 default you forgot about.

Want Your Site Legible to Every Engine — Traditional or Generative?

GEO rewards the same foundation we build for search: machine-readable structure, clean crawl paths, and content with something original in it.

About the Author

Shakur Abdirahman
Technical SEO Specialist
Shakur is a Technical SEO Specialist who implements directly inside WordPress and Shopify environments. He helps small and large sites become legible to every engine — traditional search and generative assistants alike — through machine-readable structure, clean crawl paths, and content built to be cited.

More about the author →