Skip to content

25.9%

13.3%Anthropic and Common Crawl

Share of high-value training tokens restricted to OpenAI crawlers

How much of the good web has been closed to AI training?

25.9% of tokens in the head of the C4 corpus were restricted to OpenAI's crawlers by April 2024, against 13.3% for Anthropic and Common Crawl.

ValueApril 2024 · Global
  • 25.9%

    Share of high-value training tokens restricted to OpenAI crawlers

    April 2024 · Global

  • 13.3%

    Anthropic and Common Crawl

    For comparison

The only non-commercial measurement of this, and the clearest evidence that the opt-out tokens changed behaviour: restrictions redistributed sharply right after GPTBot and Google-Extended were introduced. It is also the oldest figure in this group, and the direction has almost certainly continued since.

Source

Reference period
Geography
Global
Evidence type
Analysis
Sample
Tokens in Head C4, the 3,950 highest-token domains of the C4 corpus. Audit of 14,000 web domains with longitudinal robots.txt collection
For comparison
Anthropic and Common Crawl: 13.3%
Last checked

Method

Data Provenance Initiative, "Consent in Crisis", audit of 14,000 web domains underlying C4, RefinedWeb and Dolma, with longitudinal robots.txt collection from archived snapshots and human annotation. Head C4 is the 3,950 highest-token domains. Data ends April 2024, so treat it as the historical inflection rather than current state. The paper's forward projections are forecasts and are not used here.

Method →

Related questions

Related indicators