Frequently asked question · Optimize Content for AI
Should we block AI crawlers to protect our content?
You can, and it costs you the citation without protecting much. An engine that cannot fetch your page still answers the question, from a Reddit thread, a review site, or a competitor’s comparison, and describes you from those. The audit tells GPTBot, which OpenAI uses for training, apart from OAI-SearchBot, which ChatGPT uses to fetch a page it is about to cite, and the same for the other engines, so you can allow the retrieval crawlers and decide about training separately. What we will not recommend is a site-wide disallow you inherited from a CDN default.
What blocking actually changes
The engines answer the question whether or not they can read you. When an engine cannot fetch your page, it does not leave you out of the answer; it builds the answer from the pages it can fetch, which for most brands means a Reddit thread, a review site, a directory, or a competitor's comparison table. Your name still appears. The description of you is just sourced from people who are not you. That is the trade: a blocked crawler protects the page and gives up any say in what is said about it.
Training and retrieval are separate crawlers
The crawlers have different jobs, and robots.txt lets you treat them differently. GPTBot is what OpenAI uses to collect training data. OAI-SearchBot is what ChatGPT uses to fetch a page it is about to cite in a live answer. The other engines split the same way, and Google-Extended, PerplexityBot, and ClaudeBot are among the crawlers the audit checks. The GEO audit checks robots.txt access for 20 AI crawlers on one URL per template and reports each as pass, warn, or fail, with the line that caused it. So you can allow the retrieval crawlers, keep training as a separate policy decision, and see exactly which rule is doing what.
The rule nobody chose
Most site-wide blocks we find were never a decision. They came in with a CDN bot list, a WAF default, or a robots.txt line added in 2023 when the crawlers were new and the answer engines did not exist. The audit surfaces that line, and on Managed GEO we change it and run the same URL again once the fix is live. On GEO Consulting and Trends the fix goes to your platform team per template, and we re-audit when they say it has shipped.
Answered on
Optimize Content for AI
Make your content retrievable and citable by every AI engine: crawler access, answer-first pages, the sources engines lean on, and citation rate per engine.
More questions about Optimize Content for AI
See what AI is already saying about your brand.
A free audit shows you exactly where you stand across ChatGPT, Gemini, Perplexity, Copilot, Claude, Grok, and AI Overviews, and what it would take to close the gap.