Check which AI crawlers can access your website
Enter your domain to see which AI crawlers your robots.txt allows or blocks, including GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and Google-Extended. Search bots decide whether your pages can be used in AI answers; training bots only decide whether your content can be used to train models. Each result shows the robots.txt line behind the verdict and a snippet you can copy to change it.
Which AI crawlers exist, and what does blocking each kind do?
AI companies run different crawlers for different jobs, and each one has its own robots.txt name. Blocking the wrong one is the usual mistake: it can keep your pages out of AI answers, or it can have no effect at all.
Search bots build the index behind answers
OAI-SearchBot (ChatGPT search), Claude-SearchBot, PerplexityBot, DuckAssistBot and Amzn-SearchBot find pages to show and link in AI search. Googlebot and Bingbot build the Google and Bing indexes that Google's AI Overviews, Google's AI Mode and Microsoft's AI answers draw on. Blocking a search bot makes your pages unlikely to appear in that product.
User-triggered fetchers read one page on request
ChatGPT-User, Claude-User, Perplexity-User, meta-externalfetcher and Amzn-User load a page because a person asked an assistant about it. Anthropic documents that Claude-User follows robots.txt; OpenAI, Perplexity, Meta and Amazon say theirs may not.
Training crawlers feed future models
GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, meta-externalagent, Amazonbot, CCBot and Bytespider collect public pages for model training or control whether they may be used for it. Blocking them is a business decision, not an SEO requirement, and it does not remove you from AI search. Google-Extended and Applebot-Extended are control tokens that do not crawl.
Fix it: a robots.txt you can copy
Choose what you want, copy the result and add it to the robots.txt at the root of your site. Keep your existing rules, and edit an existing group for the same bot instead of adding a second one.
# AI search bots: allowed
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
User-agent: DuckAssistBot
User-agent: Amzn-SearchBot
Allow: /
# User-triggered fetchers: allowed
User-agent: ChatGPT-User
User-agent: Claude-User
User-agent: Perplexity-User
User-agent: meta-externalfetcher
User-agent: Amzn-User
Allow: /
- If a "User-agent: *" rule blocks everything, the allowed groups still let these bots in. Googlebot and Bingbot stay under your normal rules.
- A bot with its own group ignores the "User-agent: *" group completely, so copy any path rules (for example, Disallow: /wp-admin/) into its group if they should still apply to that bot.
- Crawlers re-read robots.txt only from time to time, so a change can take a day or more to apply. Blocking training does not remove you from AI search, but blocking search bots lowers your chance of being used in AI answers.
- If a bot is allowed here but never seems to visit, check the bot settings of your CDN, hosting firewall and security plugins: some block AI crawlers by default.
The crawlers we check
The checker covers 21 crawlers from AI and search companies. Roles and robots.txt behavior follow each vendor's documentation, last checked on 2 October 2026.
| Crawler (robots.txt name) | Vendor | Role | Follows robots.txt? | Vendor documentation |
|---|---|---|---|---|
Googlebot |
Search index | Yes, per the vendor | Documentation | |
Bingbot |
Microsoft | Search index | Yes, per the vendor | Documentation |
Applebot |
Apple | Search index | Yes, per the vendor | Documentation |
OAI-SearchBot |
OpenAI | Search index | Yes, per the vendor | Documentation |
Claude-SearchBot |
Anthropic | Search index | Yes, per the vendor | Documentation |
PerplexityBot |
Perplexity | Search index | Yes, per the vendor | Documentation |
DuckAssistBot |
DuckDuckGo | Search index | Yes, per the vendor | Documentation |
Amzn-SearchBot |
Amazon | Search index | Yes, per the vendor | Documentation |
ChatGPT-User |
OpenAI | User-triggered fetch | May ignore it (user-triggered) | Documentation |
Claude-User |
Anthropic | User-triggered fetch | Yes, per the vendor | Documentation |
Perplexity-User |
Perplexity | User-triggered fetch | May ignore it (user-triggered) | Documentation |
meta-externalfetcher |
Meta | User-triggered fetch | May ignore it (user-triggered) | Documentation |
Amzn-User |
Amazon | User-triggered fetch | May ignore it (user-triggered) | Documentation |
GPTBot |
OpenAI | Model training | Yes, per the vendor | Documentation |
ClaudeBot |
Anthropic | Model training | Yes, per the vendor | Documentation |
Google-Extended |
Model training | Yes, per the vendor | Documentation | |
Applebot-Extended |
Apple | Model training | Yes, per the vendor | Documentation |
meta-externalagent |
Meta | Model training | Yes, per the vendor | Documentation |
Amazonbot |
Amazon | Model training | Yes, per the vendor | Documentation |
CCBot |
Common Crawl | Model training | Yes, per the vendor | Documentation |
Bytespider |
ByteDance | Model training | Disputed | None published |
What this check can and cannot tell you
- It reads the rules in robots.txt as our checker received them. A site can serve different files to different visitors, and a firewall or CDN can block a crawler before it ever reads robots.txt.
- robots.txt is a request, not enforcement. Reputable crawlers follow it; user-triggered fetchers may not, and some crawlers are reported to ignore it.
- We do not measure whether any AI product actually crawls, cites or shows your site. Allowed means readiness, not visibility.
- Roles and rules follow each vendor's own documentation, last checked on 2 October 2026. Vendors add and rename crawlers, so review your file at least twice a year.
Questions about AI crawlers and robots.txt
Should I block GPTBot?
It depends on whether you want OpenAI to use your public pages for model training. GPTBot is OpenAI's training crawler, and blocking it opts your pages out of future training. That is a business decision, not an SEO requirement, and it does not affect Google rankings. If you are happy for your content to help models know your brand, you can leave it allowed.
Does blocking GPTBot remove me from ChatGPT search?
No. ChatGPT search uses a different crawler, OAI-SearchBot, and pages a user asks ChatGPT to open are fetched by ChatGPT-User. Blocking GPTBot only affects training. Blocking OAI-SearchBot is what makes your pages unlikely to appear in ChatGPT search results.
What is the difference between GPTBot and OAI-SearchBot?
GPTBot collects public pages that may be used to train OpenAI's generative models. OAI-SearchBot finds pages to show and link in ChatGPT's search features. They are separate robots.txt names, so you can allow one and block the other.
Which crawlers must I allow to appear in AI answers?
Allow the search bots: OAI-SearchBot, Claude-SearchBot, PerplexityBot, DuckAssistBot, and Googlebot and Bingbot, which build the indexes that Google's and Microsoft's AI answers draw on. Allowing them makes your pages eligible, not guaranteed: robots.txt only decides whether a bot may crawl, not whether anyone cites you.
Is robots.txt enough to keep AI crawlers out?
No. It is a request that reputable crawlers follow, not enforcement. User-triggered fetchers can ignore it, and some crawlers are reported to. To block a bot more reliably, use a firewall or CDN rule. Remember that blocking a search bot also keeps you out of that product.
Does blocking Google-Extended hide me from Google AI Overviews?
No. Google-Extended is a control token for Gemini training and grounding. AI Overviews and AI Mode use the normal Google Search index crawled by Googlebot, and Google says Google-Extended does not affect Search inclusion or ranking. To limit how your content appears there, use snippet controls such as nosnippet or max-snippet, which also affect regular snippets.
Want the full picture?
The free analysis checks robots.txt together with structured data, speed, SEO and AI citability. If you want someone to fix what it finds, GE-KO builds and optimizes websites.