Who blocks AI crawlers?
We checked whether robots.txt disallows the homepage for 15 AI crawlers, grouped by purpose: 8 training crawlers (they collect pages to train models), 4 AI search crawlers (they build an index an assistant can cite from) and 3 user-triggered fetchers (they load a page because a person asked an assistant about it). Two of the training entries, Google-Extended and Applebot-Extended, are control tokens rather than crawlers. The guide to AI crawlers and robots.txt explains the difference.
Show the data as a table
| Check | Sites | Out of | Share (%) |
|---|---|---|---|
| Blocks at least one AI crawler | 18 | 90 | 20 |
| Blocks a training crawler | 18 | 90 | 20 |
| Blocks an AI search crawler | 14 | 90 | 16 |
| Blocks a user-triggered fetcher | 11 | 90 | 12 |
| Blocks all AI crawlers | 1 | 90 | 1 |
Show the data as a table
| Check | Sites | Out of | Share (%) |
|---|---|---|---|
| Bytespider | 17 | 90 | 19 |
| ClaudeBot | 16 | 90 | 18 |
| CCBot | 14 | 90 | 16 |
| GPTBot | 14 | 90 | 16 |
| Applebot-Extended | 12 | 90 | 13 |
| Google-Extended | 11 | 90 | 12 |
| Amazonbot | 10 | 90 | 11 |
| meta-externalagent | 9 | 90 | 10 |
| PerplexityBot | 12 | 90 | 13 |
| DuckAssistBot | 8 | 90 | 9 |
| Claude-SearchBot | 5 | 90 | 6 |
| OAI-SearchBot | 5 | 90 | 6 |
| ChatGPT-User | 8 | 90 | 9 |
| Perplexity-User | 6 | 90 | 7 |
| Claude-User | 5 | 90 | 6 |
The share of sites that block a given crawler ranges from 6% (for example OAI-SearchBot) to 19% (Bytespider). Googlebot is blocked by 0 of 90 sites.
Blocking a training crawler is a legitimate choice and does not remove a site from AI search. Blocking an AI search crawler can keep a site out of the index an assistant cites from, and 16% of the sites block at least one. We did not measure whether any of these sites appears in AI answers.
On 2 October 2026, we fetched robots.txt again for the 18 sites flagged as blocking and read it with a second, simpler parser. The blocked crawlers matched for 18 of 18. We did not recheck sites that were not flagged.
Readiness signals beyond crawler access
We also looked for these readiness signals: an llms.txt file (a proposed Markdown summary of a site; no major search engine has committed to using it, see the llms.txt guide), a sitemap declared in robots.txt, schema.org structured data and hreflang annotations. None of them guarantees that an AI answer will cite a site.
Show the data as a table
| Check | Sites | Out of | Share (%) |
|---|---|---|---|
| Has an llms.txt file | 12 | 90 | 13 |
| Declares a sitemap in robots.txt | 70 | 84 | 83 |
| Homepage has schema.org markup | 50 | 85 | 59 |
| Homepage has Organization markup | 30 | 85 | 35 |
| Schema.org markup on most pages | 55 | 83 | 66 |
| FAQPage markup on at least one page | 8 | 83 | 10 |
| Language annotations (hreflang) on at least one page | 31 | 83 | 37 |
We counted structured data by presence and did not validate it (see the structured data guide). FAQ markup is optional, so a missing FAQPage is not an error by itself. The hreflang annotation matters only for multilingual sites.
On 2 October 2026, we fetched the 12 detected llms.txt files again. Of these, 12 are real text files, not HTML error pages.
On-page basics that search and AI both read
Search engines and AI systems both read basic page elements such as the title, description, headings and image alt text. The charts show how many sites have gaps in them, using the definitions of the Sitequiry Site Report (see the methodology).
Show the data as a table
| Check | Sites | Out of | Share (%) |
|---|---|---|---|
| Meta description missing on at least one page | 47 | 81 | 58 |
| Meta description missing on most pages | 15 | 81 | 19 |
| H1 heading missing on more than one page | 35 | 81 | 43 |
| Open Graph tags incomplete on at least one page | 41 | 81 | 51 |
| Images without alt text on at least one page | 40 | 83 | 48 |
| More than a tenth of images without alt text | 13 | 83 | 16 |
Across 83 sites, we checked 71,406 images. 8,991 of them (13%) have no alt attribute at all. An empty alt="" is correct for decorative images, so we did not count it.
By language group
We show a figure for a language group only if it is based on at least 15 sites; otherwise the cell reads "n too small". The language is the main language of the site as we classified it. The sample has 36 Slovenian, 26 German (mainly Germany and Austria), 19 English and 25 sites in other languages (Croatian, Italian, French, Spanish, Czech, Dutch). These counts include sites that could not be analyzed.
Differences between groups can say more about which sites we picked than about a language or a country, and the groups are small. Read the table as a description, not as a ranking.
The English-language group
The English-language group has 15 completed reports. Of these sites, 7% block at least one AI crawler (1 of 15; whole sample: 20%) and 20% publish llms.txt (3 of 15; whole sample: 13%).
Page-level figures cover only 14 English-language sites, fewer than the 15 we require, so we show them for the whole sample only.
| Check | All sites | Slovenian | German | English | Other languages |
|---|---|---|---|---|---|
| Blocks at least one AI crawler | 20% · 18/90 | 15% · 5/33 | 30% · 7/23 | 7% · 1/15 | 26% · 5/19 |
| Blocks an AI search crawler | 16% · 14/90 | 12% · 4/33 | 17% · 4/23 | 7% · 1/15 | 26% · 5/19 |
| Has an llms.txt file | 13% · 12/90 | 12% · 4/33 | 17% · 4/23 | 20% · 3/15 | 5% · 1/19 |
| Homepage has schema.org markup | 59% · 50/85 | 61% · 20/33 | 70% · 16/23 | n too small (n = 13) | 63% · 10/16 |
| Homepage has Organization markup | 35% · 30/85 | 42% · 14/33 | 39% · 9/23 | n too small (n = 13) | 25% · 4/16 |
| Meta description missing on at least one page | 58% · 47/81 | 74% · 23/31 | 43% · 10/23 | n too small (n = 13) | n too small (n = 14) |
| H1 heading missing on more than one page | 43% · 35/81 | 48% · 15/31 | 22% · 5/23 | n too small (n = 13) | n too small (n = 14) |
| Open Graph tags incomplete on at least one page | 51% · 41/81 | 48% · 15/31 | 57% · 13/23 | n too small (n = 13) | n too small (n = 14) |
| Images without alt text on at least one page | 48% · 40/83 | 56% · 18/32 | 43% · 10/23 | n too small (n = 14) | n too small (n = 14) |
Methodology
This is a convenience sample, not a random one. We chose 106 well-known public websites to cover different types and languages: 31 shops and marketplaces, 22 news and magazine sites, 25 public bodies, health and education sites, 15 companies, SaaS and agencies, 8 tourism and hospitality sites and 5 non-profits. Of these sites, 10 are known to use bot protection.
- Crawl. On 29 September 2026, SitequiryBot, the crawler of the Sitequiry Site Report, read each site: the homepage plus further pages, up to 20 pages in total, picked from the sitemap and the main navigation. These pages are not a random sample of a site. The crawler respects robots.txt and pauses between requests. Software: Site Report build 98996b7.
- Completed reports. We completed 90 of 106 reports. We left out 13 sites that blocked the first request (for example with HTTP 403 or 429) and 3 that redirected to another domain.
- Blocking AI crawlers. We count a crawler as blocked when robots.txt disallows the homepage ("/") for it, either in a group that names the crawler or in the "*" group when the crawler has no group of its own. Matching follows RFC 9309. A missing or unreadable robots.txt counts as "no rules", and rules for sections of a site are not counted.
- Pages. Page-level figures describe a site only when at least 5 of its pages could be analyzed. Of the completed reports, 7 had fewer analyzable pages, because of client-rendered pages without content in the HTML, bot protection after the first request, or very little crawlable content. A URL that redirects to another page is not counted twice. We judge title, description, H1 and Open Graph on indexable pages only (no noindex, canonical pointing to the page itself). We do not judge markup and headings on pages whose HTML holds no content because a script renders it.
- Images. An image counts as missing alt text only when its img element has no alt attribute. We ignore tracking pixels and hidden images.
- Shares. Every share is the number of sites with the property divided by the number of sites in the population named in the dataset. The dataset lists both numbers for every figure.
Limits of this study
- Not representative. We chose the sites ourselves, mostly well-known organizations. The study says little about small business websites.
- Sites that blocked our request are not in the figures, so the results lean toward sites that allow automated checks.
- robots.txt is a request, not enforcement. A site may also block crawlers with a firewall or CDN, and may serve different rules to different crawlers. We read robots.txt as SitequiryBot.
- Readiness signals, not visibility. We did not measure whether a site is cited or shown in AI answers, and the presence of llms.txt or markup is not proof of any effect.
- One crawl from one location on 29 September 2026. Sites change, and we read up to 20 pages per site.
- Small groups. Language groups are small, and a single site moves a share by several points.
- No causal claims. The data shows what sites do, not why.
What this study does not measure
Each item below is left out for the reason given.
- Speed. Google PageSpeed was not run for this sample, so there are no Core Web Vitals, no distribution of mobile scores and no render-blocking or page-weight figures.
- Server response time. It was recorded for single requests from one location only; it is not a speed measurement.
- HTTPS redirects and security headers. The crawl did not record them.
- Visibility in AI answers. Only readiness signals are measured.
- Validity of structured data. It is counted by presence.
- Quality of llms.txt files. They are counted by presence.
- Sitequiry scores. The SEO, AEO, GEO and overall scores are our own weighted rubric, not an independent measure.
- Results per site. Only aggregates are published.
- Segments by site type or country. They would have fewer than 15 sites each.
Cite this study and download the data
You may quote the figures and reuse the charts with attribution. The aggregate data is licensed under Creative Commons Attribution (CC BY 4.0): you may share and adapt it, including commercially, if you credit GE-KO / Sitequiry and link to this page.
GE-KO / Sitequiry (2026). How ready are websites for AI search? A readiness-signal analysis of 90 public websites. Published 2 October 2026. https://sitequiry.com/studies/ai-search-readiness-2026
Version 1.0, licensed CC BY 4.0. Column names and labels in the files are in English.
Summary for press and newsletters
Sitequiry, the free website analysis tool by the Slovenian agency GE-KO, analyzed 90 well-known public websites in a crawl on 29 September 2026. 20% block at least one AI crawler in robots.txt, 13% publish an llms.txt file and 59% of homepages carry schema.org markup. The authors chose the sites, which are not representative of the web, and the study measures readiness signals, not visibility in AI answers.
What to do next
- Decide deliberately which AI crawlers you allow, and check your firewall and CDN too: AI crawlers and robots.txt.
- Add an llms.txt only if it is cheap to keep up to date: llms.txt: what it is and whether you need one.
- Describe who you are and what a page is about with accurate markup: structured data guide.
- Check your own site against these signals: free analysis, one URL, full result, no account.
- Want the fixes done? GE-KO implements them.
Frequently asked questions
How many websites block GPTBot?
In our sample, 14 of 90 sites (16%) block GPTBot, OpenAI's training crawler, in robots.txt. ClaudeBot is blocked by 18% of sites, PerplexityBot by 13% and Google-Extended by 12%. ChatGPT search uses a different crawler, OAI-SearchBot, which 6% block.
What is llms.txt and how many sites have it?
llms.txt is a proposed plain Markdown file at the root of a website that summarizes the site for language models. 12 of 90 sites (13%) publish one. It is not a web standard and no major search engine has committed to using it; see the llms.txt guide.
How many websites use structured data?
In our sample, 59% of homepages carry schema.org markup (JSON-LD, microdata or RDFa), and 78% of sites carry it on at least one analyzed page. We counted presence, not validity. See the structured data guide.
Does blocking AI crawlers keep a site out of AI answers?
It depends on which crawler is blocked. Blocking a training crawler does not remove a site from AI search, while blocking an AI search crawler can keep it out of the index an assistant cites from. This study cannot show the effect, because it reads robots.txt rules, not AI answers. See the guide to AI crawlers and robots.txt.
How reliable are these numbers?
The counts are exact for the 90 sites we analyzed on 29 September 2026, but they do not describe the web as a whole. We chose the sites ourselves, 16 sites could not be analyzed, and the language groups are small. The methodology and limits are above, and the aggregate data is open under CC BY 4.0.
Check your own site in about two minutes
The free analysis reads your robots.txt, llms.txt and structured data and shows what to fix first.