Playbook · GEO
How to get cited by generative engines
A model does not rank you. It reads a handful of sources and writes an answer, then names some of them. This is about being one it names.
Short answer
To get cited by generative engines, give a model something specific it can attribute: a figure with a source, a quotable sentence from someone with standing, and a clear statement of who is speaking. Peer-reviewed work on this found optimisation of that kind lifts visibility by up to 40%, and that keyword stuffing — the classic SEO reflex — performs about 10% worse than doing nothing.
The checks, in order of consequence
- 01
Give it a number worth quoting
Concrete figures, with the source of the figure named. A model reaching for something attributable will take a sentence carrying a statistic over a sentence carrying an adjective.
Statistics addition is one of the strategies the GEO paper found effective across domains.
- 02
Include a quotable line from someone with standing
A direct quotation, attributed to a named person with a reason to be listened to. Models quote quotes — the attribution is already built in.
Quotation addition was the paper's strongest method in several domains, including explanation-type queries.
- 03
Cite your own sources visibly
Link out to what you relied on. Counter-intuitive for classic SEO, which hoards link equity, but a page that shows its working reads as more citable to a model assessing it.
- 04
Stop keyword stuffing
The one tactic the research found actively harmful. It is also the one most likely to survive in an old content brief.
The paper reports keyword stuffing performing about 10% worse than the unoptimised baseline.
- 05
Write fluently
Readable, well-formed prose outperforms dense keyword-led writing. Optimising for a machine reader and writing well are, for once, the same instruction.
- 06
Cover the whole subject, not one page of it
Models retrieve from sources they can rely on across a topic. One strong page inside an otherwise thin site rarely becomes the source of record.
This is topical authority doing the work — depth is what makes a source retrievable.
- 07
Be present where models retrieve from
Coverage on third-party sites, directories and publications a model already reads. Presence in the retrieval set is upstream of everything on your own domain.
- 08
State who is speaking
Author, credentials, organisation — in text and in schema. A model deciding whom to name needs someone to name.
- 09
Let the AI crawlers in
GPTBot, PerplexityBot, ClaudeBot, Google-Extended. Blocking them is a legitimate choice, but it is a choice to be absent, and it is frequently made by accident.
The free robots.txt check tells you which are allowed on any domain.
- 10
Expect it to vary by domain
The same tactic does not pay equally everywhere. The research is explicit that effectiveness differs by subject area, so treat any universal ranking of tactics with suspicion.
What the research actually found
The best-known study here is peer-reviewed, and worth separating from the summaries of it that circulate. Figures widely quoted as per-method gains — "+27.8% for quotations", "+25.9% for statistics" — do not appear in the paper. These do.
| Finding | What the paper states |
|---|---|
| Headline result | Optimisation of this kind can boost visibility by up to 40% in generative engine responses. |
| Benchmark | GEO-bench: 10,000 queries drawn from 25 domains, spanning simple to multi-faceted questions. |
| Keyword stuffing | Performs about 10% worse than the unoptimised baseline. |
| Effect varies | Efficacy differs across domains — the paper calls for domain-specific optimisation rather than one universal recipe. |
Figures below are quoted from the paper itself, not from summaries of it.
Source
Aggarwal et al., “GEO: Generative Engine Optimization”, KDD ’24 (arXiv:2311.09735)
Why citation is not ranking
In classic search, ten results are shown and you compete for one of the slots. A generative engine reads a small set of sources and writes a single answer, naming some of what it read. You are not competing for a position — you are competing to be inside the set it drew from, and then to be worth naming.
That changes what you optimise. Retrievability comes first: can a model find and rely on you across the subject, not just on one page. GEO is the name for that work, and topical authority is the mechanism behind it.
Read the research, not the summaries of it
Most articles on this subject cite the same paper and attach precise per-method percentages to it. Those numbers are not in the paper. What is in it is a headline of up to 40%, a 10,000-query benchmark, a clear finding that keyword stuffing is counterproductive, and an explicit warning that effects vary by domain.
The distinction matters more than it looks. A number invented in a summary and repeated confidently is exactly the kind of thing a generative engine will pick up and repeat back — which is a reason to be careful about what you publish, not only about what you read.
Questions people ask about this
Is GEO just SEO with a new name?
It shares most of its foundation — you cannot be cited by a model that cannot crawl or index you. What differs is the target. SEO optimises for a position; GEO optimises for being retrieved and named inside someone else's answer.
Can you guarantee a citation in ChatGPT?
No, and nobody can. Generative answers vary between sessions for the same question. What you can do is raise the probability, and measure it by asking the same questions repeatedly over time rather than celebrating a single good answer.
Should I block AI crawlers to protect my content?
It is a real trade-off, not an obvious call. Blocking protects content from being trained on and summarised without a click; it also removes you from the sources those engines can cite. Which side wins depends on whether your traffic converts on the page or your brand needs to be named.
How do I measure GEO at all?
Ask a fixed set of questions across several engines on a schedule and record whether you are named. It is a sampling exercise, not a rank tracker — the same question can produce different sources on different days.
Want this run against your own site?
Send the URL and I'll come back with which of these checks the site passes today, and which one is costing you the most.