~80%
72%a year earlier
Share of AI crawling that is for training rather than retrieval
Are AI crawlers collecting training data or answering questions?
Training accounted for close to 80% of AI bot activity by July 2025, up from 72% a year earlier.
~80%Share of AI crawling that is for training rather than retrieval
Worth separating from the visibility question, because they pull in opposite directions. A site that blocks training crawlers is not thereby making itself more visible in answers, and a site that wants to appear in answers is not obliged to hand over training data. The two controls are separate at every major publisher.
Source
- Reference period
- Geography
- Global
- Evidence type
- Platform telemetry
- Sample
- Share of AI bot activity Cloudflare classifies as training rather than search or user action. Dominated by one bot, OpenAI's GPTBot
- For comparison
- a year earlier: 72%
- Last checked
We have not opened the primary instrument for this figure. It comes either from the publisher's own summary or from a document that cites it, so treat the level as indicative.
Method
Cloudflare, to roughly July 2025. The split rests on Cloudflare's own classification of each user agent's declared purpose, the taxonomy carries an undeclared bucket, and the aggregate is dominated by one bot, OpenAI's GPTBot. It is therefore closer to a statement about OpenAI than about AI crawling in general.
Related questions
Related indicators
None
Extra technical requirements to appear in Google's AI Overviews
3 of 4
1 of 4state no such exceptionAssistant makers whose docs say user-triggered fetches may ignore robots.txt
4.2%
4.5%Googlebot aloneAI bots as a share of HTML requests
~14%
312domains disallowing GPTBotDomains with AI-specific robots.txt directives
25.9%
13.3%Anthropic and Common CrawlShare of high-value training tokens restricted to OpenAI crawlers
55.3%