Data Study · AI SEO

    1.5 Million Websites Block at Least One AI Crawler From Reading Their Pages

    Matt SuffolettoWritten by Matt Suffoletto
    Published Sept. 24, 2026 5 min read
    Share

    1.5 million websites, 9.9% of the 15.5 million in the public HTTP Archive crawl, block at least one AI crawler. They block the bots that collect training material about four times as often as the bots that fetch pages for live AI answers.

    Analysis of 15,470,478 websites · the public HTTP Archive crawl, September 2026

    Key Findings

    1. 1.1.5 million websites, 9.9% of the web, block at least one AI crawler from reading their pages.
    2. 2.OpenAI's training crawler, GPTBot, is the most-blocked AI bot, turned away by 8.7% of websites.
    3. 3.Only 2.1% of websites block the OpenAI bots that fetch pages for ChatGPT's search answers, and 1.7% block PerplexityBot.
    4. 4.28.9% of the most visited 1,000 websites block at least one AI crawler, nearly three times the 9.9% rate across all websites.
    5. 5.More than 1.1 million websites block each of eight different AI training crawlers, from GPTBot to Meta's crawler.

    Summary

    Every website can post a short file, robots.txt, that tells automated visitors which pages they may read. Since AI companies began collecting the web to train their models and answer questions, a long list of AI crawlers has appeared in those files, each with an instruction to stay out.

    For business owners, that one file decides two different things. Blocking a training crawler keeps your content out of the material future AI models learn from. Blocking a search bot can keep your pages out of the live answers ChatGPT and Perplexity give, and so out of their citations. Some site owners never chose what their file says, because a host, theme or plugin wrote it for them.

    1.5 million websites, 9.9% of the 15.5 million in the HTTP Archive, block at least one AI crawler. Nine in ten still let every AI bot in. Those that block mostly target training crawlers: 8.7% of websites block GPTBot, against 2.1% that block ChatGPT's search bots.

    What we measured

    We read the robots.txt file of every homepage in the HTTP Archive's September 2026 mobile crawl, a public record of 15,470,478 websites, and recorded which AI crawlers each file blocks. We report the share of all websites, and the share of the most visited 1,000 sites as ranked by Google's Chrome UX Report.

    A crawler, or bot, is a program that visits pages automatically. We checked 14 AI crawlers: training crawlers such as OpenAI's GPTBot, Common Crawl's CCBot, Anthropic's ClaudeBot, ByteDance's Bytespider, Amazonbot, Google-Extended, Applebot-Extended and Meta-ExternalAgent, and the bots behind live AI answers, OpenAI's OAI-SearchBot and ChatGPT-User and PerplexityBot. A site blocks a bot when its robots.txt has a rule addressed to that bot by name; a rule aimed at all bots at once does not count.

    One website in ten blocks AI crawlers

    9.9%
    of websites block at least one AI crawler.

    That is 1,532,161 websites. Among sites with a working robots.txt file, the rate is 11.5%. Nine in ten websites still let every AI crawler read every page, which means the great majority of the web remains open to both AI training and AI search.

    In absolute terms, that is a large group. Any business checking whether AI engines can read its pages should start with its own robots.txt file, because the site may be among the 1.5 million without the owner knowing.

    Training crawlers are blocked four times as often as search bots

    8.7%
    of websites block GPTBot, OpenAI's training crawler, more than any other AI bot.

    ClaudeBot (8.3%), CCBot (8.2%) and ByteDance's Bytespider (7.9%) follow closely. The bots that fetch pages for live answers sit at the bottom: 2.1% of websites block OpenAI's search agents and 1.7% block PerplexityBot. GPTBot is blocked about four times as often as ChatGPT's search bots.

    Share of all websites blocking it by AI crawler
    GPTBot
    8.7%
    ClaudeBot / anthropic-ai / Claude-Web
    8.3%
    CCBot
    8.2%
    Bytespider
    7.9%
    Amazonbot
    7.8%
    Google-Extended
    7.6%
    Applebot-Extended
    7.2%
    Meta-ExternalAgent
    7.1%
    OAI-SearchBot or ChatGPT-User
    2.1%
    PerplexityBot
    1.7%

    Source: Suff Digital analysis of 15,470,478 websites · the public HTTP Archive crawl, September 2026

    This split matters. Blocking GPTBot keeps a site out of OpenAI's future training data, but it does not remove the site from ChatGPT's search answers, which rely on the separate search bots. Most blocking websites keep training out and leave search open.

    Most blockers block many bots at once

    The counts for individual training crawlers are tightly bunched. 1.53 million websites block at least one AI bot, and more than 1.1 million block each of the eight training crawlers. Most blocking sites turn away many AI agents in one go, the pattern of a shared block list rather than a bot-by-bot decision.

    Share of all websites by AI crawler
    GPTBot
    8.7%
    ClaudeBot / anthropic-ai / Claude-Web
    8.3%
    CCBot
    8.2%
    Bytespider
    7.9%
    Amazonbot
    7.8%
    Google-Extended
    7.6%
    Applebot-Extended
    7.2%
    Meta-ExternalAgent
    7.1%
    OAI-SearchBot or ChatGPT-User
    2.1%
    PerplexityBot
    1.7%
    Any AI bot
    9.9%

    Source: Suff Digital analysis of 15,470,478 websites · the public HTTP Archive crawl, September 2026

    AI crawler Websites blocking Share of all websites Share of the top 1,000
    GPTBot 1,349,614 8.7% 23.1%
    ClaudeBot / anthropic-ai / Claude-Web 1,280,614 8.3% 21.6%
    CCBot 1,272,870 8.2% 21.6%
    Bytespider 1,218,191 7.9% 19.5%
    Amazonbot 1,207,354 7.8% 15.6%
    Google-Extended 1,178,614 7.6% 19.2%
    Applebot-Extended 1,120,241 7.2% 13.5%
    Meta-ExternalAgent 1,100,573 7.1% 15.6%
    OAI-SearchBot or ChatGPT-User 321,615 2.1% 12.6%
    PerplexityBot 258,945 1.7% 14.1%
    Any AI bot 1,532,161 9.9% 28.9%

    The most visited sites block three times as often

    28.9% of the most visited 1,000 websites block at least one AI crawler, nearly three times the 9.9% rate across the web. They block GPTBot at 23.1%, and they are far more willing to cut off live AI search too: 14.1% block PerplexityBot and 12.6% block ChatGPT's search bots, against 1.7% and 2.1% of all websites.

    Large sites have content, audiences and licensing options that make blocking a deliberate business decision. For most smaller sites, the calculation runs the other way, because being read and cited by AI engines is a source of new visitors.

    What this means for website owners

    Before you add AI bots to robots.txt, decide what you want. To stay out of model training, block GPTBot, Google-Extended, CCBot and ClaudeBot. To appear in AI answers and get cited, leave OAI-SearchBot, ChatGPT-User and PerplexityBot open. Most small business sites benefit from being cited, so blanket blocking rarely makes sense for them.

    Then check what your site does today. Some hosts, themes and plugins add AI blocks without the owner deciding, and an old rule can quietly shut out the bots that bring AI search visitors. A robots.txt and crawler-access review is a standard part of technical seo services, and it takes minutes to fix once found.

    Embed this research

    Paste this on your site to embed the charts. It links back to the source automatically.

    <iframe id="sd-ai-crawler-blocking-rate" src="https://www.suffdigital.com/embed/data-studies/ai-crawler-blocking-rate" width="100%" height="600" style="width:100%;border:1px solid #E5E7EB;border-radius:12px" loading="lazy" title="1.5 Million Websites Block at Least One AI Crawler From Reading Their Pages - Suff Digital"></iframe>
    <script>window.addEventListener("message",function(e){if(e&&e.data&&e.data.sdEmbed==="ai-crawler-blocking-rate"&&e.data.height){var f=document.getElementById("sd-ai-crawler-blocking-rate");if(f){f.style.height=e.data.height+"px";}}});</script>
    <p style="font:14px/1.5 system-ui,sans-serif">Source: <a href="https://www.suffdigital.com/resources/data-studies/ai-crawler-blocking-rate">1.5 Million Websites Block at Least One AI Crawler From Reading Their Pages - Suff Digital</a></p>

    Cite this study

    APA

    Suff Digital. (2026). 1.5 Million Websites Block at Least One AI Crawler From Reading Their Pages. https://www.suffdigital.com/resources/data-studies/ai-crawler-blocking-rate

    Plain link

    1.5 Million Websites Block at Least One AI Crawler From Reading Their Pages - Suff Digital - https://www.suffdigital.com/resources/data-studies/ai-crawler-blocking-rate

    Frequently asked questions

    Related studies