Getting cited by ChatGPT and Perplexity: the GEO guide
get cited by ChatGPT generative engine optimization GEO llms.txt guide AI search optimization

Getting cited by ChatGPT and Perplexity: the GEO guide

6 min read
Back to articles
TL;DR Quick summary for those in a hurry

Your future clients don't always google anymore: they ask ChatGPT, Perplexity or Google's AI Overviews. Getting cited by these engines — GEO, Generative Engine Optimization — comes down to concrete levers: allowing AI crawlers in your robots.txt, publishing an llms.txt, structuring every page as extractable answers, and offering dated, sourced figures that models love to quote. A practical guide for independents and small businesses, tested on this very site.

Ask ChatGPT "which real-estate photographer would you recommend on the Costa Daurada?" or "how do I choose an SEO consultant?". Somebody gets cited. The question that matters for your business: why not you? This guide isn't theoretical — every technique described here runs on the site you're reading, and part of my traffic already arrives from visitors sent by ChatGPT and Perplexity. Here's the method, from server to writing style.

First, understand how an AI "chooses" whom to cite

Two distinct mechanisms, two workstreams:

  1. Training: models remember what they read during learning. Slow, diffuse, outside your direct control.
  2. RAG (live retrieval): ChatGPT Search, Perplexity and AI Overviews query the web in real time and cite sources. This is where your game is played — and the criteria look like classic SEO pushed to its logical extreme: accessible, structured, factual, trustworthy content.

Practical consequence: GEO doesn't replace your SEO, it rewards it. A site invisible on Google will be invisible to Perplexity too. If your foundations are shaky, start there before optimising for AI.

Workstream 1: open the technical doors

  • Allow AI crawlers in your robots.txt

    GPTBot (OpenAI), ClaudeBot (Anthropic), PerplexityBot, Google-Extended: each is controlled individually. Blocking them 'on principle' removes you from the conversation — OpenAI documents its bots and user-agents at platform.openai.com/docs/bots.

  • Publish an llms.txt file

    An emerging standard (spec at llmstxt.org): a Markdown file at your site root presenting your key pages to models. Ten minutes of work — this site has one at amory-studio.com/llms.txt if you want a live example.

  • Check your site renders without JavaScript

    Several AI crawlers read raw HTML. A site whose content only appears after JavaScript execution is unreadable to them — test with 'view source', not the inspector.

  • Keep structured data clean

    Schema.org markup (Article, FAQPage, LocalBusiness…) helps engines understand who you are and what you claim — the same markup that powers your Google rich results. Reference: schema.org.

The technical references for this workstream: the llms.txt specification, the OpenAI bots documentation and the Schema.org vocabulary.

Workstream 2: write to be quoted

Generative engines assemble answers from fragments. Your goal: make every section of every page a self-contained, quotable fragment.

Practice Invisible-to-AI version Quotable version
Structure Long paragraphs that 'flow' H2 phrased as the question + complete 40-80 word answer right below
Data 'Professional photos speed up sales' 'Listings with professional photos sell X% faster (source, year)'
Positioning 'We are digital leaders' 'Real-estate photographer in Tarragona, HDR specialist, since 2019'
Freshness Undated article Visible update date + genuinely refreshed content
ℹ️
The 10-second test :

Take any page on your site and ask: if an AI extracts ONLY the first paragraph under each heading, is the answer complete, factual, dated? If yes, you're quotable. That's exactly what the TL;DR boxes and FAQs on this site are for.

Workstream 3: become the source AIs prefer

Models preferentially cite what looks like a primary source: original data, lived experience, identifiable expertise. Three accessible levers for an independent:

  1. Publish your own numbers: your real prices, average turnaround times, a mini-comparison of your local market. A "what does X cost in 2026" page with real figures becomes THE extractable reference of your niche.
  2. Sign and embody: name, bio, concrete experience on every piece (E-E-A-T). Google's helpful-content guidance — documented on Google Search Central — also shapes what AIs pick up.
  3. Keep your local footprint consistent: AIs cross-reference your site with your listings and directories. An inconsistent NAP blurs your entity — see our guide on NAP and local citations.

Measure: AI traffic is already visible

In your analytics (Umami, Plausible, GA4), build a referrer segment for chatgpt.com, perplexity.ai, claude.ai and copilot.microsoft.com. Two tracking habits: the monthly trend of those visits, and a quarterly manual test — ask ChatGPT and Perplexity the 5 questions your clients ask you, and note who gets cited. That's your next-generation ranking, and it moves fast. Site speed matters in the crawl equation too — see our take on what the PageSpeed score really means.

FAQ

What is GEO (Generative Engine Optimization)?
The set of practices that increase your chances of being cited as a source by generative engines — ChatGPT Search, Perplexity, Google's AI Overviews. It combines classic SEO foundations with specific levers: allowed AI crawlers, llms.txt, extractable answers, dated and sourced data.
Should I block or allow GPTBot and AI crawlers?
For a business seeking clients, allowing them is usually right: blocked, your content disappears from the answers your prospects read. Each bot (GPTBot, ClaudeBot, PerplexityBot, Google-Extended) is controlled individually in robots.txt.
What is an llms.txt file?
A Markdown file at the site root (emerging standard documented at llmstxt.org) presenting your essential pages to AI models. Quick to create, it acts as a sitemap designed for LLMs — this site publishes one at /llms.txt.
Does GEO replace classic SEO?
No — it builds on it and rewards it: generative engines rely heavily on the same signals (accessibility, structure, authority, freshness). A site that ranks poorly on Google has very little chance of being cited by Perplexity or ChatGPT Search.
How do I measure traffic coming from AIs?
Create a referrer segment for chatgpt.com, perplexity.ai, claude.ai and copilot.microsoft.com in your analytics tool, and add a quarterly manual test: ask the AIs your clients' typical questions and note whom they cite.

Your business deserves to be cited by AIs

I run GEO on my own site — and I can audit yours: SEO foundations, llms.txt, quotable content and AI-traffic tracking.

Request an audit