SEO + GEO · 9 min read
How Perplexity, ChatGPT, and Claude Choose What to Cite

AI citations — the sources that Perplexity, ChatGPT, and Claude surface when they answer a question — are not random. Each engine applies a consistent set of signals to decide which pages earn a reference. Understanding those signals is now a core competency for any marketing director who wants their brand to appear in AI-generated answers.
Why AI citations are the new first page
When a buyer asks ChatGPT “which CRM is best for a 50-person sales team,” the answer cites two or three sources. Those sources get the click, the brand impression, and the implied endorsement. Everything else is invisible. AI citations have replaced the first page of Google results as the primary attention surface for a growing share of commercial queries — and that share is accelerating.
Marketing directors who still measure success purely by organic rank are optimising for a surface that is shrinking. The question is no longer “how do I rank?” It is “how do I get cited?” The mechanics are different, and most teams have not caught up.
How each engine retrieves sources
The three dominant AI answer engines use different retrieval architectures, and those differences matter for strategy.
Perplexity: live web retrieval
Perplexity runs a live web search before generating its answer. It fetches pages in real time, parses their content, and selects sources based on relevance, freshness, and structural clarity. A page that loads slowly, blocks crawlers, or buries its key claim in dense prose is at a structural disadvantage. If you want to go deeper on Perplexity-specific tactics, the Perplexity optimization playbook covers the retrieval layer in detail.
ChatGPT: browsing plus training weight
ChatGPT with browsing enabled behaves similarly to Perplexity for recent queries. For evergreen topics, it also draws on training data, which means older, authoritative pages can earn AI citations without being freshly crawled. The implication: domain authority still matters, but it works through a different mechanism than traditional SEO. The ChatGPT citations playbook breaks down how to influence both the training-weight and the live-retrieval layers.
Claude: document and context grounding
Claude is most often used with documents or context injected by the operator. When it does retrieve from the web, it prioritises sources with clear entity definitions, explicit authorship, and structured claims. Claude is particularly sensitive to whether a page answers a question directly rather than building to an answer over several paragraphs.
The five signals that drive AI citations
Across all three engines, five signals consistently separate cited pages from uncited ones. These are not hypotheses — they are observable patterns from pages that earn AI citations repeatedly across different query types.
- Topical authority: The domain publishes consistently on a narrow subject. A page about CRM on a site that covers CRM deeply earns AI citations more often than the same page on a generalist site.
- Direct answer placement: The page answers the query in the first 60–100 words. AI engines parse for the answer, not the build-up.
- Entity clarity: The page names the subject, defines it, and connects it to related entities explicitly. Vague language (“our solution helps teams”) does not register as a named entity.
- Structured markup: Pages with schema — FAQ, HowTo, Article — give the engine a machine-readable map of the content. The schema markup guide for AI snippets explains which types carry the most weight for AI citations.
- Crawlability and freshness: A page the engine cannot access does not get cited. Neither does a page whose last-modified date signals it has not been touched in three years on a fast-moving topic.
Entity clarity: the signal most sites miss
Most marketing content is written for humans who already know the brand. It uses pronouns, shorthand, and assumed context. AI engines do not share that context. They need explicit entity signals: the company name, the product category, the geography, the use case, and the relationship between them — stated plainly, not implied.
A page that says “we help mid-market companies close deals faster” gives an AI engine almost nothing to work with. A page that says “Acme CRM is a sales pipeline tool built for B2B companies with 20–200 sales reps, headquartered in Austin, Texas” gives it five distinct entities to anchor AI citations. This is why brand mentions now outperform backlinks for GEO — a mention that includes entity context is worth more than a link that provides none.
How to build entity signals into existing content
You do not need to rewrite everything. Audit your top 20 pages and check whether each one states the subject entity, the category, and the use case in the first two paragraphs. If it does not, add a single dense sentence that does. That one change, applied consistently, is often enough to shift a page from uncited to cited within a few crawl cycles.
Content structure and AI citations
AI engines parse structure before they parse prose. A page with clear headings, short paragraphs, and explicit answers to sub-questions is easier to cite than a long-form essay that buries its claims. This is not about dumbing down content — it is about making the content’s architecture legible to a machine that has milliseconds to decide whether to use it.
The content cluster model for AI search visibility is the structural approach that works best here. A pillar page that defines the topic, supported by cluster pages that answer specific sub-questions, gives AI engines a clear hierarchy to navigate. Each cluster page can earn AI citations independently, and the pillar page benefits from the topical authority the cluster builds.
FAQ sections as citation targets
FAQ sections are disproportionately cited because they are pre-formatted as question-answer pairs — exactly the structure an AI engine is trying to produce. Every substantive page on your site should end with three to five FAQ items that address the real questions buyers ask. Mark them up with FAQ schema. Google’s structured data documentation explains the technical implementation, and the same markup signals apply to AI retrieval engines.
Technical foundations that affect AI citations
Content quality is necessary but not sufficient. A technically broken page does not get cited regardless of how good the writing is. Three technical factors have an outsized effect on AI citations.
- Crawl access: If your robots.txt blocks AI crawlers — Perplexity’s bot is PerplexityBot, OpenAI’s is GPTBot — you are invisible. Check your crawl configuration before anything else.
- Page speed: Perplexity fetches pages in real time. A page that takes four seconds to load is a page that may time out before it is parsed. Site speed is a GEO signal, not just an SEO one — the mechanics are different but the outcome is the same for AI citations.
- Clean HTML: AI parsers prefer clean, semantic HTML. Pages built on heavy JavaScript frameworks that render content client-side are harder to parse than pages that serve content in the initial HTML response. This is a build decision with real citation consequences.
Before and after: a GEO-optimised page
| Signal | Typical unoptimised page | GEO-optimised page |
|---|---|---|
| Opening paragraph | Brand story, mission statement | Direct answer to the target query in <80 words |
| Entity definition | Implied (“our platform”) | Explicit (product name, category, use case, geography) |
| Heading structure | Marketing copy (“Why choose us?”) | Query-matched (“How does X compare to Y?”) |
| Schema markup | None or basic Article | Article + FAQ + BreadcrumbList |
| Crawl access | Unchecked; may block AI bots | Verified open for PerplexityBot, GPTBot, ClaudeBot |
| Page speed | 3–5 second load | Sub-1.5 second LCP, clean server-rendered HTML |
What to prioritise first
If you are starting from zero, the sequence matters. Doing everything at once produces nothing measurable. Here is the order that generates the fastest signal for earning AI citations:
- Week 1: Audit crawl access. Confirm AI bots can reach your top 30 pages. Fix any blocks.
- Week 2–3: Add direct-answer opening paragraphs to your top 20 pages. One dense, entity-rich paragraph per page.
- Week 4–6: Add FAQ sections with schema markup to those same pages. Three to five questions per page, matched to real buyer queries.
- Week 7–10: Build or restructure your content cluster. Identify the pillar topic, map the sub-questions, and publish or update cluster pages to answer each one directly.
- Ongoing: Monitor which pages earn AI citations using Perplexity’s citation tracker and manual spot-checks. Iterate on the pages that are close but not yet cited.
AI citations compound. A page that earns one citation gets crawled more frequently, which increases the chance of earning more. The early movers in any category are building a lead that will be hard to close later.
If you want to map this against your specific content library and competitive landscape, the ChatGPT citations playbook is a good next read — or talk to Studio Máté about running a GEO audit on your site.
FAQ
What is the single most important factor for earning AI citations?
Direct answer placement. AI engines parse for the answer first. A page that states its key claim in the opening paragraph — clearly, with named entities — is cited far more often than a page that builds to the same claim over several sections. Everything else amplifies this, but nothing replaces it.
Do backlinks still matter for AI citations?
They matter indirectly. Backlinks contribute to domain authority, which influences training-data weight in models like ChatGPT. But for live-retrieval engines like Perplexity, the page-level signals — structure, entity clarity, crawlability — carry more weight than the link graph. A well-structured page on a mid-authority domain will often outperform a thin page on a high-authority domain.
How do I know if my pages are being cited by AI engines?
Run the queries your buyers actually ask in Perplexity, ChatGPT, and Claude. Check whether your domain appears in the citations. Do this for 20–30 queries across your core topics. The pattern will tell you which pages are working and which are invisible. There is no automated tool that covers all three engines reliably yet — manual spot-checking is still the most accurate method.
Does schema markup directly cause AI citations?
Schema does not cause citations, but it removes friction. FAQ and Article schema give the engine a machine-readable map of the content, which makes it easier to parse and cite. Pages with schema consistently outperform equivalent pages without it, all else being equal. The effect is strongest for FAQ schema on pages that already have strong entity signals and direct-answer structure.
How often should I update pages to maintain AI citations?
For evergreen topics, a quarterly review is sufficient — check that the facts are current and the entity signals are still accurate. For fast-moving topics (pricing, product comparisons, regulatory changes), update within days of a material change. Perplexity in particular weights freshness heavily for queries where recency is implied by the question.