Where the term comes from
The term was introduced in a paper titled "GEO: Generative Engine Optimization", submitted to arXiv on 16 November 2023 by Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan and Ameet Deshpande, and last revised on 28 June 2024.
This matters more than it usually would, because GEO is a term with a paper trail. Most marketing acronyms do not have one, and the paper says several things that later usage has quietly dropped.
The authors' framing is about a power imbalance rather than a marketing tactic. They describe content creators as "the third stakeholder", who, "given the black-box and fast-moving nature of generative engines", have "little to no control over when and how their content is displayed." GEO is presented as a response to that: a "flexible black-box optimization framework" for a system nobody outside it can see into.
What the paper actually measured
The headline result is the one that travels: the paper demonstrates that GEO "can boost visibility by up to 40% in generative engine responses."
Three things belong next to that number, and all three are in the same abstract.
"Up to" is a ceiling, not an average. It describes the best observed case, not what a given site should expect.
Effectiveness is domain-dependent. The authors state that "the efficacy of these strategies varies across domains, underscoring the need for domain-specific optimization methods." A method that worked in one field is not established for another.
It was measured on their own benchmark. The evaluation ran on GEO-bench, which the same paper introduces: "a large-scale benchmark of diverse user queries across multiple domains, along with relevant web sources to answer these queries." That is a reasonable way to make results reproducible, and it is not the same as a measurement in the live systems people actually use.
The paper dates from 2023 and 2024. The generative engines it studied have changed since.
The number that travels, and the three things the same abstract says next to it.
up to 40%
in generative engine responses
“can boost visibility by up to 40% in generative engine responses”
arXiv 2311.09735, abstract
A ceiling, not an average
what “up to” means
It describes the best observed case, not what a given site should expect.
The paper reports a maximum. A maximum quoted as a forecast is a different claim than the one measured.
It depends on the domain
not transferable as-is
A method that worked in one field is not established for another.
“the efficacy of these strategies varies across domains, underscoring the need for domain-specific optimization methods”
Measured on GEO-bench
their own benchmark
A reasonable way to make results reproducible — and not the same as a measurement in the systems people actually use.
“a large-scale benchmark of diverse user queries across multiple domains, along with relevant web sources to answer these queries”
What Google says about its own AI features
Google's documentation for site owners takes a different position on its own surfaces, and states it plainly:
"There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary."
And, specifically on the file-and-markup approaches often sold under the GEO label:
"You don't need to create new machine readable files, AI text files, or markup to appear in these features. There's also no special schema.org structured data that you need to add."
This is not a contradiction of the paper, and reading it as one gets both wrong. The two statements have different scopes. The paper is about generative engines as a class, including assistants that are not Google. Google is describing what is required to appear in [AI Overviews](/glossary/ai-overview) and AI Mode specifically.
What follows is narrower and more useful than either "GEO works" or "GEO is snake oil": for Google's AI surfaces, the operator of those surfaces says the existing fundamentals are the mechanism. For assistants that are not Google, that statement does not apply, and no equivalent statement from those vendors is quoted here because none was found.
Two statements that sound opposed. They are about different things.
The paper
arXiv 2311.09735
“can boost visibility by up to 40% in generative engine responses”
Applies to: generative engines as a class, including assistants that are not Google. Measured 2023–2024.
Search Central documentation
“There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary.”
Applies to: AI Overviews and AI Mode specifically — the operator describing its own surfaces.
What that leaves worth doing
The fundamentals Google names are unglamorous and apply regardless: pages that can be crawled, content findable through internal links, important information available as text, structured data that matches what a reader sees.
Beyond that, the honest position is that visibility in AI answers is easier to measure than to optimise. Measuring is a matter of asking assistants what they say about a category and recording the answer. Optimising is a matter of claims, and the strongest published claim comes with "up to", a domain caveat and a benchmark from 2024.
Anyone selling a GEO service should be able to say which of those two they are doing.
Common mistakes
Quoting the 40 percent without its conditions. It is an upper bound from a specific benchmark, measured on systems that have since changed.
Treating GEO as a Google technique. Google says its AI features need no special optimisation. GEO as a term covers generative engines broadly, most of which are not Google.
Buying AI-specific files or markup for Google visibility. The documentation addresses this directly and says they are not needed. Whether they help elsewhere is a separate question with a separate answer.
Assuming what worked in one sector transfers. The paper's own finding is that effectiveness varies by domain.
Related terms
- [AI Overview](/glossary/ai-overview): the Google surface most GEO discussion is aimed at
- AI Mode: a separate Google surface with different models and different links
- llms.txt: one of the file-based approaches Google's documentation addresses
- Answer engine: the broader category the paper calls a generative engine
Source
- Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, Deshpande, "GEO: Generative Engine Optimization", arXiv:2311.09735, submitted 16 November 2023, last revised 28 June 2024: https://arxiv.org/abs/2311.09735 (retrieved 26 August 2026). Source of the term's origin, the "third stakeholder" framing, the black-box description, the "up to 40%" result, the domain-variance caveat and the description of GEO-bench. All quotations are from the abstract as published on that page.
- Google Search Central, "AI features and your website": https://developers.google.com/search/docs/appearance/ai-features (retrieved 26 August 2026; page states last updated 2025-12-10). Source of both quoted statements on additional requirements, files, markup and structured data.
Frequently asked questions
Is GEO just a new name for SEO?
No. SEO optimizes for ranking position in a list of links; GEO optimizes for being cited or mentioned inside a generated answer. They share a technical foundation, since a page still needs to be crawlable and indexed for either, but the writing practices that help most differ between the two.
Do I need to abandon my SEO strategy to do GEO?
No, and doing so would likely hurt both. Generative AI systems draw on the same indexed web content traditional search relies on, so a page still needs solid technical SEO to be eligible for citation. GEO adds writing and structuring practices on top of that foundation rather than replacing it.
How do I know if GEO is working?
Run the actual prompts your buyers would ask across the AI assistants they use, repeated over time, and check whether your domain gets cited or mentioned more often than before. A single check does not show a trend; the same prompt set needs to be run consistently to see real movement.
Does GEO work the same way across ChatGPT, Perplexity and Google's AI Overviews?
Not identically. Each system retrieves from different sources, weighs different signals, and updates on a different schedule, so a content change that increases citations on one system will not automatically produce the same result on another. Testing needs to happen per system, not once for all of them.
