training crawlers
CCBot
Operated by Common Crawl · Training crawlers
Common Crawl's bot. It archives public pages into a free, open dataset that many model builders later train on.
When CCBot shows up in your logs
Your page is being saved into Common Crawl's open dataset, which feeds many downstream models.
Crawlers that collect public content that may be used to train or ground models.
Allow or block it in robots.txt
A robots.txt entry is a request. Well-behaved crawlers honor it; it is not an enforcement mechanism.
Is CCBot on your site right now?
Voris AI visibility answers that: verified AI agent traffic, split from your human visitors, on every plan. Cookieless, with a free plan and a 14-day trial that needs no card.