Official documentation, crawler user agents, specifications, research, and tools for Answer Engine Optimization (AEO) and Generative Engine Optimization (GEO) in 2026.
Answer Engine Optimization (AEO) and Generative Engine Optimization (GEO) are the practice of making web content discoverable, retrievable, and citable by answer engines such as ChatGPT, Google AI Overviews and AI Mode, Perplexity, Google Gemini, Claude, Microsoft Copilot, and Grok. The two terms are used interchangeably by most practitioners. Where a distinction is drawn, Answer Engine Optimization (AEO) emphasizes being cited in a direct answer, and Generative Engine Optimization (GEO) emphasizes influencing the generated text itself.
- Tools
- Official Engine Documentation
- Crawler User Agents
- Specifications and Standards
- Research Papers and Datasets
- Analytics and Measurement
- aeo-radar - Open source. Answer Engine Optimization monitor for tracking brand visibility across answer engines.
- Ahrefs Brand Radar - AI visibility measurement across six AI tools, built on Ahrefs' search-backed prompt data.
- ansvisor - Open source. Tracks citations, prompts, competitors, and content opportunities; self-hosted or managed.
- aperture - Open source. AI visibility monitoring and analytics for tracking how a brand appears in answer engines.
- AthenaHQ - Answer Engine Optimization (AEO) and Generative Engine Optimization (GEO) platform for commercial and enterprise brands.
- Authoritas - SEO platform with AI search visibility tracking alongside classic rank tracking.
- Botify - AI search optimization platform focused on large-site crawling and indexing.
- Brandlight - AI visibility platform aimed at enterprise brands.
- BrightEdge - Enterprise SEO and AI search platform covering Google Search, AI Overviews, and ChatGPT.
- canonry - Open source. Self-hosted Answer Engine Optimization (AEO) stack for tracking ChatGPT, Claude, Gemini, and Perplexity.
- Cloudflare AI Crawl Control - Monitoring and control of how AI services access a site, at the network edge rather than by robots.txt convention.
- Conductor - Enterprise Answer Engine Optimization (AEO) and SEO intelligence with website monitoring and agents.
- daydream - Full-service organic search combining SEO agents with human experts.
- Elmo - Open source. Tracks how answer engines mention, cite, and describe a brand. Self-hostable, with a hosted commercial plan from $29 per month. The #1 open source Profound replacement.
- Evertune - AI brand monitoring focused on how models represent a brand across the customer journey.
- gego - Open source. Generative Engine Optimization (GEO) tracking for a brand across multiple large language models.
- geo-aeo-tracker - Open source. Local-first AI visibility dashboard tracking a brand across six AI models.
- geo-lint - Open source. Linter applying Generative Engine Optimization (GEO), SEO, and content quality rules to pages.
- geolook - Open source. End-to-end Generative Engine Optimization (GEO) implementation covering analysis, diagnosis, and strategy.
- GEORank - Open source. Generative Engine Optimization (GEO) ranking and optimization platform.
- GetCito - Open source. Brand and competitor benchmarking across multiple AI answer surfaces.
- Goodie - AI search visibility and Answer Engine Optimization (AEO) platform for monitoring and optimizing brand presence.
- HubSpot AI Search Grader - Free one-time check of how ChatGPT, Perplexity, and Gemini describe a brand.
- Knowatoa - AI search visibility tracking oriented toward recovering traffic lost to answer engines.
- Known Agents - Directory and analytics for AI agents and bots, formerly Dark Visitors.
- Known Agents Directory - Continuously updated catalog of AI crawler user agents and their operators.
- LLMrefs - Brand visibility, rank, and citation tracking across generative answer engines.
- Nightwatch - Rank tracker unifying classic search positions with AI visibility in ChatGPT, Claude, Gemini, and Perplexity.
- oneglanse - Open source. Free Generative Engine Optimization (GEO) tracker for monitoring brand appearance in answer engines.
- Otterly.AI - AI search monitoring for ChatGPT, Perplexity, and Google AI Overviews.
- Peec AI - AI search analytics for marketing teams, benchmarking brand performance against competitors.
- Profound - Brand visibility measurement and optimization for answer engines.
- Rankscale - AI visibility and ranking tracker across ChatGPT, Perplexity, Gemini, and Google AI Overviews.
- Relixir - Generative Engine Optimization (GEO) monitoring paired with automated content generation and deployment.
- Scrunch AI - AI search visibility monitoring, site optimization, and content delivery to AI agents.
- SE Ranking AI Visibility Tool - Brand mention and link tracking in AI answers, with competitor comparison.
- searchstack-aeo - Open source. Answer Engine Optimization (AEO), Generative Engine Optimization (GEO), and SEO stack aimed at small teams.
- Semrush Enterprise - Enterprise SEO and AI search platform.
- seoClarity - Unified SEO and Answer Engine Optimization (AEO) platform for enterprise teams.
- Similarweb - Digital market intelligence, including traffic measurement for AI assistant referrals.
- Superlines - AI search intelligence for brands and agencies.
- Trakkr - Citation, perception, and competitor tracking across ChatGPT, Claude, and Gemini.
- XFunnel - Citation tracking and question discovery across AI search platforms.
- Yext Scout - AI search visibility agent scanning multiple models with competitor comparison.
- ZipTie.dev - Tracker for Google AI Overviews, ChatGPT, and Perplexity.
Vendor-published documentation on crawling, citations, publisher controls, and attribution. Every entry is hosted on a domain the vendor itself controls.
- Overview of OpenAI Crawlers - Official reference for every OpenAI crawler, with full user-agent strings, per-bot purposes, and links to IP range files.
- GPTBot IP Ranges - Machine-readable IP prefixes for the training crawler, for verifying that a request claiming to be GPTBot really is.
- OAI-SearchBot IP Ranges - Machine-readable IP prefixes for the crawler that builds the ChatGPT search index.
- ChatGPT-User IP Ranges - Machine-readable IP prefixes for user-initiated fetches made from ChatGPT.
- OAI-AdsBot IP Ranges - Machine-readable IP prefixes for the crawler that validates advertiser landing pages.
- Publishers and Developers FAQ - OpenAI's answers on how publisher content is surfaced, cited, and controlled.
- ChatGPT Search - How ChatGPT decides to search the web and how inline citations are presented to users.
- Introducing ChatGPT Search - Launch announcement describing the citation and attribution model for publishers.
- Web Search Tool - API documentation for the web search tool, including the citation annotation format applications must display.
- AI Features and Your Website - Google's statement of how AI Overviews and AI Mode source content, and exactly which preview controls apply to them.
- Optimizing for Generative AI Features on Google Search - Google's own optimization guide, including its rebuttals of common Answer Engine Optimization (AEO) and Generative Engine Optimization (GEO) claims.
- A New Resource for Optimizing for Generative AI in Google Search - Search Central announcement introducing that guide.
- Top Ways to Ensure Your Content Performs Well in Google's AI Experiences - Earlier Search Central guidance that the optimization guide builds on.
- Google Crawlers and Fetchers Overview - Index of every Google crawler, fetcher, and robots.txt product token.
- Google's Common Crawlers - Full user-agent strings for Googlebot, GoogleOther, Google-CloudVertexBot, and the Google-Extended token.
- Google's Special-Case Crawlers - Crawlers that operate for specific products and ignore the global robots.txt user-agent rules.
- Search Generative AI Control - Search Console setting that opts a site out of generative AI features in Google Search.
- New Opportunities, Control and Insights for Website Owners - Google's announcement of that opt-out control and the reporting that accompanies it.
- AI Mode in Google Search - Product announcement describing what AI Mode is and how it links out.
- An Update on Web Publisher Controls - Google's introduction of Google-Extended and a statement of which products it governs.
- Grounding with Google Search - How Gemini grounds answers in Google Search results and returns grounding metadata and citations.
- Grounding Overview - The grounding sources available to Gemini and how each attributes its sources.
- GroundingMetadata Reference - Exact response schema for grounding chunks and supports, which is what a citation is made of.
- Perplexity Crawlers - Official reference for PerplexityBot and Perplexity-User, with full user-agent strings and IP range files.
- PerplexityBot IP Ranges - Machine-readable IP prefixes for the indexing crawler.
- Perplexity-User IP Ranges - Machine-readable IP prefixes for user-initiated fetches.
- Introducing the Perplexity Publishers' Program - The revenue-share and analytics program for publishers whose content is cited.
- Understanding Source Labels - How Perplexity labels and ranks the sources it cites in an answer.
- Does Anthropic Crawl Data From the Web, and How Can Site Owners Block the Crawler? - The single official source for ClaudeBot, Claude-User, and Claude-SearchBot, and what blocking each one does.
- Which Crawlers Does Bing Use? - Bing's own list of the crawlers behind Bing search and Microsoft Copilot.
- Announcing User-Agent Change for Bing Crawler Bingbot - The announcement that carries the current bingbot user-agent strings verbatim.
- Announcing New Options for Webmasters to Control Usage of Their Content in Bing Chat - The
NOCACHEandNOARCHIVEcontrols that govern whether Copilot may quote and link a page. - Bing Introduces Support for the data-nosnippet HTML Attribute - Element-level control over which parts of a page Bing may show.
- Introducing AI Performance in Bing Webmaster Tools - Launch announcement, and the definition of grounding queries and citations as Microsoft measures them.
- New AI Visibility Insights in Bing Webmaster Tools - The expansion that added intents, topics, citation share, and competitive comparison.
xAI publishes no crawler documentation. There is no vendor page listing Grok's user agents, no robots.txt guidance, and no published IP ranges, so this list has no entry to give you. Third-party crawler directories publish conflicting strings for xAI; because none of them is the operator, none is cited here. If xAI publishes documentation, open an issue and it will be added.
- xAI Documentation - xAI's developer documentation, listed so you can confirm for yourself that it covers the API and not crawling or publisher controls.
Not answer engines in their own right, but they crawl for AI systems and show up in the same access logs.
- About Applebot - Applebot user-agent formats plus Applebot-Extended, Apple's robots.txt opt-out for generative model training.
- Amazonbot - Amazonbot's user-agent string and Amazon's statement that crawled data may train Amazon AI models.
- Mistral Crawlers - One of the few vendors that separates training, indexing, and user-triggered fetching into three distinct user agents.
- DuckAssistBot - DuckDuckGo's real-time fetcher for AI-assisted answers, with an explicit statement that it does not train models.
- Meta Web Crawlers - Meta's list of crawler user agents, separating the training and indexing crawler from the user-request fetcher.
- CCBot - The Common Crawl crawler, whose archives are an input to many model training pipelines.
Every string below is quoted from the operator's own documentation linked in the section above. Where a vendor publishes only a robots.txt token and not a full user-agent string, that is stated rather than filled in from a third-party directory. Version numbers change without notice, so match on the token, never on the whole string.
The distinction that matters for Answer Engine Optimization (AEO) is the one between crawlers that build a corpus and fetchers that act for a user in real time. Blocking the first affects whether a model knows about you at all. Blocking the second affects whether you can be cited in an answer being generated right now. Several vendors state that user-triggered fetchers do not follow robots.txt.
Collect content that may be used to train foundation models.
| Operator | robots.txt token | Full user-agent string as documented |
|---|---|---|
| OpenAI | GPTBot |
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot |
| Anthropic | ClaudeBot |
Not published |
| Apple | Applebot-Extended |
Control token only; see below |
| Amazon | Amazonbot |
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amazonbot/0.1) Chrome/W.X.Y.Z Safari/537.36 |
| Mistral | MistralAI-Training |
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-Training/1.0; +https://docs.mistral.ai/robots) |
| Meta | meta-externalagent |
meta-externalagent/1.1 (+/documentation/sharing/webmasters/web-crawlers) |
| Common Crawl | CCBot |
CCBot/2.0 |
Build the retrieval index an answer engine draws on. These are the crawlers that determine whether you are eligible to be cited at all.
| Operator | robots.txt token | Full user-agent string as documented |
|---|---|---|
| OpenAI | OAI-SearchBot |
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot |
| Perplexity | PerplexityBot |
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot) |
| Anthropic | Claude-SearchBot |
Not published |
Googlebot |
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/W.X.Y.Z Safari/537.36 |
|
| Microsoft | bingbot |
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/W.X.Y.Z Safari/537.36 Edg/W.X.Y.Z |
| Mistral | MistralAI-Index |
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-Index/1.0; +https://docs.mistral.ai/robots) |
Google's position is that "AI is built into Search and integral to how Search functions, which is why robots.txt directives for Googlebot is the control for site owners to manage access to how their sites are crawled for Search." There is no separate AI Overviews or AI Mode crawler to allow or block. Perplexity states that PerplexityBot is used to surface and link sites in Perplexity search results and is not used to crawl content for AI foundation models.
Fetch a page in real time because a user asked a question right now. Several operators document that these fetchers do not honor robots.txt, because the request originates from a person rather than from an automated crawl.
| Operator | robots.txt token | Full user-agent string as documented |
|---|---|---|
| OpenAI | ChatGPT-User |
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot |
| Perplexity | Perplexity-User |
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user) |
| Anthropic | Claude-User |
Not published |
| Mistral | MistralAI-User |
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-User/1.0; +https://docs.mistral.ai/robots) |
| DuckDuckGo | DuckAssistBot |
DuckAssistBot/1.2; (+http://duckduckgo.com/duckassistbot.html) |
| Meta | meta-externalfetcher |
meta-externalfetcher/1.1 (+/documentation/sharing/webmasters/web-crawlers) |
Perplexity documents that Perplexity-User generally ignores robots.txt rules because the fetch is user-initiated. Meta documents that meta-externalfetcher may bypass robots.txt for the same reason.
These appear in robots.txt but never in an access log. They are opt-out switches, not user agents, and blocking them has no effect on search crawling.
| Token | Operator | What disallowing it does, per the vendor |
|---|---|---|
Google-Extended |
Excludes the site from helping improve Gemini Apps and Vertex AI generative APIs. Google documents it as a standalone product token that uses existing Google user-agent strings. | |
Applebot-Extended |
Apple | Excludes the site's content from training Apple foundation models, without affecting Applebot's search crawling. |
Google-Extended does not control AI Overviews or AI Mode. Google's AI features documentation points to nosnippet, data-nosnippet, max-snippet, and noindex for those, and Search Console carries a separate opt-out for generative AI features in Search.
User-agent strings are trivially spoofed, so treat the string as a claim and verify it.
- Verify Google Crawler Requests - Reverse-DNS and IP-range verification for Google's crawlers and fetchers.
- How to Verify Bingbot - Reverse-DNS verification for Microsoft's crawler.
OpenAI and Perplexity both take the IP-range approach instead: each publishes a JSON file of the prefixes its bots crawl from, linked in their sections above. Match the requesting address against that file rather than trusting the user-agent header.
- RFC 9309: Robots Exclusion Protocol - The robots.txt standard itself, which every AI crawler control is layered on top of.
- How Google Interprets the robots.txt Specification - The most detailed public description of real-world robots.txt parsing, including token matching and precedence.
- Robots Meta Tags Specifications - The snippet-level controls that determine how much of a page an answer engine may quote.
- IETF AI Preferences Working Group - The standards effort to replace today's incompatible per-vendor tokens with one vocabulary.
- A Vocabulary for Expressing AI Usage Preferences - The draft vocabulary of preference terms.
- Attaching AI Preferences to Content - The companion draft for expressing those preferences in robots.txt and HTTP headers.
- ai.robots.txt - Community-maintained robots.txt blocklist of known AI crawler tokens, updated as vendors add agents.
- The llms.txt Proposal - Jeremy Howard's proposal for a Markdown file that gives models a curated map of a site. It is a community convention, not a standard, and no engine documented in this list commits to reading it.
- llms-txt Repository - The reference implementation and specification source.
- llms.txt Directory - Directory of sites that publish an llms.txt file.
- llms.txt Site Index - A second index of published llms.txt files, useful for seeing real-world formatting in practice.
Weigh the effort against what the engines say. Google's optimization guide addresses llms.txt by name and states that you do not need to create new machine-readable files, AI text files, markup, or Markdown to appear in Google Search including its generative AI capabilities, "as Google Search itself doesn't use them." No other engine in this list documents reading llms.txt either. That does not make publishing one harmful; it does mean nobody has documented a benefit.
Google's guidance is explicit that no special structured data is required for AI features. Structured data still earns rich results in classic search, still disambiguates entities, and is still the cheapest way to state facts unambiguously, which is why these types stay on the list.
- Introduction to Structured Data Markup - How search systems consume schema.org markup and which formats are accepted.
- schema.org/Organization - Entity identity: names, logos,
sameAslinks, and the disambiguation an engine needs to know who you are. - schema.org/Article - Authorship, publication date, and publisher, the provenance fields that citation-bearing answers lean on.
- schema.org/FAQPage - Explicit question-and-answer pairs, the structure that most closely matches how an answer engine chunks a page.
- schema.org/QAPage - A single user-submitted question with answers, distinct from
FAQPage. - schema.org/HowTo - Ordered steps with tools and materials, for procedural answers.
- schema.org/Product - Product identity, offers, and reviews, the fields commercial answers are assembled from.
- schema.org/Dataset - Dataset descriptions, licensing, and distribution.
- schema.org/BreadcrumbList - Page position within a site, which supplies hierarchy an extractor would otherwise have to infer.
- IAB: Measuring Visibility in the AI Era - The industry framework for measuring brand and publisher visibility in AI-powered discovery, published August 2026. It defines a common metrics hierarchy called the four P's, distinguishes decision-grade from directional measurement, and sets disclosure requirements for vendors. It is the closest thing this field has to an agreed vocabulary.
Peer-reviewed and preprint work, oldest first. Vendor benchmarks and agency studies are excluded.
- GEO: Generative Engine Optimization - Aggarwal et al., 2023, accepted to KDD 2024. The paper that named Generative Engine Optimization (GEO) and the origin of most of the field's vocabulary.
- Evaluating Verifiability in Generative Search Engines - Liu, Zhang, and Liang, 2023. Human evaluation of whether generative search engine citations actually support the sentences they are attached to.
- CC-GSEO-Bench: A Content-Centric Benchmark for Measuring Source Influence in Generative Search Engines - Chen et al., 2025. A benchmark for measuring how much an individual source shapes a generated answer.
- Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement - Sielinski, 2026. Treats answer-engine visibility as a sampling problem, which is the right frame for anyone reading a visibility dashboard.
- From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms - Zhang, He, and Yao, 2026. Separates being cited from actually influencing the answer text, and measures both across platforms.
- Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines - Kumar, 2026. A large-scale measurement of brand visibility across engines.
- Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization, 2023-2026 - Martinez, 2026. A survey of the field to date and the best single starting point for the literature.
Where each engine's activity actually shows up. Two different things get measured and they are easy to confuse: whether you were cited inside an answer, and whether a person clicked through to your site. Vendor consoles report the first. Your own analytics report the second, and only for the fraction of answers that produce a click at all.
- Generative AI Performance Report for Search - Search Console impressions for AI Overviews and AI Mode, broken down by page, country, device, and date.
- Generative AI Performance Report for Discover - The equivalent report for generative AI features in Google Discover.
- Introducing Search Generative AI Performance Reports - Google's announcement of those reports and what the metrics do and do not include.
- Bing Webmaster Tools AI Performance - Citation counts, page-level performance, and the grounding queries Copilot generated internally to find your content.
Google reports impressions in generative AI features. Microsoft reports citations and grounding queries, and since June 2026 also intent labels, topic groups, and citation share; both Microsoft announcements are linked in the Copilot section above. Impressions and citations are not the same metric and should not be summed.
Neither Perplexity, OpenAI, Anthropic, nor xAI operates a public webmaster console reporting citations back to site owners. For those engines, referral traffic and third-party trackers are all there is.
- GA4 Default Channel Group - Google Analytics documents an
AI Assistantschannel for arrivals from sources like ChatGPT, Gemini, Copilot, and Grok, assigned when the medium isai-assistantor the referrer matches Google's list of AI assistants. Google notes this channel excludes AI Overviews and AI Mode, which are reported as ordinary organic search. - Custom Channel Groups - How to build your own grouping if you need engines split individually rather than pooled.
Referral analytics only ever sees the answers that produced a click. To see the fetches that produced none, you need the server-side view from an edge tool such as Cloudflare AI Crawl Control, listed under Tools above.
Clicks from an answer engine arrive as ordinary referral traffic with that engine's hostname as the referrer. No vendor in this list publishes a specification of its referrer strings, so any fixed list of them is an observation rather than documentation, and this list does not publish one. Read the referrer hostnames out of your own logs and confirm them against your own traffic before hard-coding them into a report.
Contributions are welcome. Read the contribution guidelines first: every entry must link to a primary source, and no user-agent string is accepted unless the operator publishes it. See also the code of conduct.