Skip to content

Question

Does blocking AI crawlers in robots.txt keep my hotel out of ChatGPT?

Short answer

It stops the scheduled crawlers. It does not reliably stop a fetch a person triggered, because three of the four major assistant makers publish that robots.txt may not apply to those.

Two different things get called the same thing. A training crawler visits on a schedule to collect content; a user-triggered fetch happens when someone asks a question and the assistant goes to look. The first obeys robots.txt at every major publisher. The second is where the documentation diverges.

OpenAI writes that because those actions are initiated by a user, robots.txt rules may not apply. Google writes that its user-triggered fetchers generally ignore robots.txt. Perplexity writes the same of Perplexity-User. Anthropic states its bots honour robots.txt and names no exception, which is not the same as promising one.

The practical consequence for a property is that a blanket block is a decision about training data, not about visibility. It will keep your content out of future model weights. It will not reliably keep your page from being read when a traveller asks about you, and at OpenAI and Perplexity the search and training controls are explicitly separate: you can allow one and refuse the other.

That separation is the part worth acting on. Allowing the search bot while disallowing the training bot is a documented, supported configuration at both, and it is closer to what most hotels actually want than either extreme.

Why credible sources differ

The four publishers describe their own behaviour and nobody audits them. Cloudflare has accused Perplexity of crawling from undeclared agents after being blocked; Perplexity says Cloudflare misattributed a third party's traffic. Neither published reproducible evidence, and both sell something in the dispute.