25.9%
13.3%Anthropic and Common Crawl
Share of high-value training tokens restricted to OpenAI crawlers
How much of the good web has been closed to AI training?
25.9% of tokens in the head of the C4 corpus were restricted to OpenAI's crawlers by April 2024, against 13.3% for Anthropic and Common Crawl.
25.9%
Share of high-value training tokens restricted to OpenAI crawlers
April 2024 · Global
13.3%
Anthropic and Common Crawl
For comparison
The only non-commercial measurement of this, and the clearest evidence that the opt-out tokens changed behaviour: restrictions redistributed sharply right after GPTBot and Google-Extended were introduced. It is also the oldest figure in this group, and the direction has almost certainly continued since.
Source
- Reference period
- Geography
- Global
- Evidence type
- Analysis
- Sample
- Tokens in Head C4, the 3,950 highest-token domains of the C4 corpus. Audit of 14,000 web domains with longitudinal robots.txt collection
- For comparison
- Anthropic and Common Crawl: 13.3%
- Last checked
Method
Data Provenance Initiative, "Consent in Crisis", audit of 14,000 web domains underlying C4, RefinedWeb and Dolma, with longitudinal robots.txt collection from archived snapshots and human annotation. Head C4 is the 3,950 highest-token domains. Data ends April 2024, so treat it as the historical inflection rather than current state. The paper's forward projections are forecasts and are not used here.
Related questions
Related indicators
42%
29%sometimes fact-checkConsumers who always fact-check AI outputs
25%
Travellers who received outdated or inaccurate AI travel information
70%
Tourism businesses already using AI
94%
60%properties with under 10 staffCybersecurity readiness gap by property size
52%
Executives reporting AI agents in production
6