training crawlers

GPTBot

Operated by OpenAI · Training crawlers

OpenAI's training crawler. It gathers public pages that may feed a future GPT model.

When GPTBot shows up in your logs

OpenAI is collecting public content that could train a later model. Robots.txt controls whether it may.

Crawlers that collect public content that may be used to train or ground models.

Allow or block it in robots.txt

A robots.txt entry is a request. Well-behaved crawlers honor it; it is not an enforcement mechanism.

allow
User-agent: GPTBot
Disallow:
block
User-agent: GPTBot
Disallow: /

Is GPTBot on your site right now?

Voris AI visibility answers that: verified AI agent traffic, split from your human visitors, on every plan. Cookieless, with a free plan and a 14-day trial that needs no card.