~14%
312domains disallowing GPTBot
Domains with AI-specific robots.txt directives
How many sites have taken a position on AI crawlers?
About 14%, or 546 of the 3,816 top domains where a robots.txt file was found. GPTBot was both the most blocked and the most explicitly allowed.
~14%Domains with AI-specific robots.txt directives
The denominator does the work here. It is 14% of domains that have a findable robots.txt, not 14% of the top ten thousand, and it counts allow directives as well as disallow. Read as "most sites have no AI rule at all", which is the finding, rather than as "14% block AI", which is not what was measured.
Source
- Reference period
- Geography
- Global
- Evidence type
- Analysis
- Sample
- 546 of 3,816 domains that had a findable robots.txt, drawn from Cloudflare Radar's top 10,000. Counts allow as well as disallow directives
- For comparison
- domains disallowing GPTBot: 312
- Last checked
We have not opened the primary instrument for this figure. It comes either from the publisher's own summary or from a document that cites it, so treat the level as indicative.
Method
Cloudflare Radar, point-in-time crawl of 6 June 2025, drawn from Radar's top 10,000 domains. 312 domains disallowed GPTBot, 250 fully and 62 partially; 61 explicitly allowed it. Cloudflare notes some robots.txt tokens, Google-Extended among them, are not user-agent substrings and so do not match request logs.
Related questions
Related indicators
42%
29%sometimes fact-checkConsumers who always fact-check AI outputs
25%
Travellers who received outdated or inaccurate AI travel information
70%
Tourism businesses already using AI
94%
60%properties with under 10 staffCybersecurity readiness gap by property size
52%
Executives reporting AI agents in production
6