free reference

The AI crawler directory

Every AI assistant, search index, and training crawler that might visit a website: what it does, whether its traffic can be verified, and how to allow or block it. Built from the providers’ own published sources.

every crawler, one directoryai answers16 crawlerssearch indexes14 crawlerstraining crawlers19 crawlersother ai bots13 crawlersyour site62 crawlers, 21 providers, four intents. reviewed july 2026.
62crawlers documented
21providers, majors to newcomers
4intents: answers, indexing, training, other
July 2026last reviewed against provider sources

Look up any crawler

Search by crawler or provider, or filter by what the bot is for. Every entry links to its own page with verification guidance, sample IP ranges where the provider publishes them, and ready-made robots.txt lines.

All 62 crawlers

ChatGPT-UserAI answers
OpenAI

OpenAI's live fetcher. When someone asks ChatGPT a question your page can answer, this is what opens it in the moment.

IP-verifiable
Claude-UserAI answers
Anthropic

Anthropic's live fetcher for Claude. It reads your page the instant a person's question calls for it.

IP-verifiable
Perplexity-UserAI answers
Perplexity

The fetch behind a Perplexity answer. It reads your page so Perplexity can quote it and list it as a source.

IP-verifiable
Google-AgentAI answers
Google

One of Google's user-triggered agents. It opens a page in the moment, acting for a person rather than fetching an answer.

IP-verifiable
Google-NotebookLMAI answers
Google

The fetcher behind Google's NotebookLM. It opens a page when someone adds it as a source to research with.

IP-verifiable
Google-Read-AloudAI answers
Google

A Google fetcher for read-aloud and assistant features. It opens your page so the text can be spoken back.

IP-verifiable
GoogleAgentAI answers
Google

Another label in Google's user-triggered agent family. Like Google-Agent, it fetches pages while carrying out a person's task.

IP-verifiable
MistralAI-UserAI answers
Mistral

Mistral's live fetcher for Le Chat. It opens your page when a person's question needs it.

IP-verifiable
CopilotAI answers
Microsoft

Microsoft Copilot fetching live. It reads your page in the moment to answer whatever someone just asked.

IP-verifiable
Amzn-UserAI answers
Amazon

Amazon's live fetcher behind products like Alexa, opening your page to pull a fresh answer for someone.

IP-verifiable
DuckAssistBotAI answers
DuckDuckGo

DuckDuckGo's real-time fetcher for DuckAssist. It reads your page so the AI answer can quote and cite it.

IP-verifiable
xAI-SearchBotAI answers
xAI

xAI's live fetcher for Grok. It opens your page as Grok answers, though xAI publishes no IP ranges to confirm it.

UA only
Grok-DeepSearchAI answers
xAI

xAI's fetcher for Grok's deep search mode. It reads your page live while building a longer, researched answer.

UA only
meta-externalfetcherAI answers
Meta

A Meta fetcher that opens a specific page when a person requests or shares its link inside a Meta product.

IP-verifiable
Kimi-UserAI answers
Moonshot AI

Moonshot's live fetcher for Kimi. It opens your page when a question there needs an answer.

IP-verifiable
Qwen-UserAI answers
Alibaba

Alibaba's live fetcher for Qwen, opening your page mid-answer. Alibaba publishes no IP ranges, so it's user-agent only.

UA only
OAI-SearchBotSearch indexes
OpenAI

OpenAI's crawler that discovers and refreshes pages so they can turn up in ChatGPT search.

IP-verifiable
Claude-SearchBotSearch indexes
Anthropic

Anthropic's crawler that indexes pages so Claude can find and cite them when it searches the web.

IP-verifiable
PerplexityBotSearch indexes
Perplexity

Perplexity's indexing crawler, quietly keeping its answer engine's map of the web current.

IP-verifiable
Google-InspectionToolSearch indexes
Google

Google's inspection fetcher, the one behind Search Console's URL inspection and the Rich Results test.

IP-verifiable
GooglebotSearch indexes
Google

Google's main search crawler, indexing the web for Search. Cloudflare notes the same traffic can feed AI Overviews.

IP-verifiable
MistralAI-IndexSearch indexes
Mistral

Mistral's indexing crawler, building the map Le Chat draws on when it answers.

IP-verifiable
BingbotSearch indexes
Microsoft

Microsoft's long-running search crawler, indexing pages for Bing and, by extension, Copilot.

IP-verifiable
msnbotSearch indexes
Microsoft

A legacy Microsoft search crawler, largely superseded by Bingbot but still turning up in logs.

IP-verifiable
Amzn-SearchBotSearch indexes
Amazon

Amazon's crawler that indexes pages to make them eligible for Amazon's search experiences.

IP-verifiable
meta-webindexerSearch indexes
Meta

Meta's indexing crawler, working to sharpen the results Meta AI returns in search.

IP-verifiable
Kimi-SearchBotSearch indexes
Moonshot AI

Moonshot's search crawler for Kimi, indexing pages so the assistant can find them when answering.

IP-verifiable
TikTokSpiderSearch indexes
ByteDance

ByteDance's crawler that indexes pages for TikTok's search and discovery surfaces.

UA only
BaiduspiderSearch indexes
Baidu

Baidu's main search crawler, indexing pages for China's largest search engine.

UA only
YouBotSearch indexes
You.com

You.com's crawler, indexing pages for its search and answer engine.

UA only
GPTBotTraining
OpenAI

OpenAI's training crawler. It gathers public pages that may feed a future GPT model.

IP-verifiable
ClaudeBotTraining
Anthropic

Anthropic's crawler for gathering public web text that may go toward improving Claude.

IP-verifiable
GoogleOtherTraining
Google

A catch-all Google crawler that product teams use to pull public pages for research and development.

IP-verifiable
Google-ExtendedTraining
Google

Not a crawler at all, but a robots.txt token that decides whether Google may use your content to train and ground Gemini.

robots.txt token
Google-CloudVertexBotTraining
Google

Google's crawler for on-demand fetches a site owner requests while building a Vertex AI agent.

IP-verifiable
Applebot-ExtendedTraining
Apple

A robots.txt token rather than a bot. It tells Apple whether it may use already-crawled content for AI training.

robots.txt token
ApplebotTraining
Apple

Apple's crawler, which Cloudflare places across both search and training uses, feeding Siri, Spotlight, and Safari.

IP-verifiable
AmazonbotTraining
Amazon

Amazon's general crawler, gathering public content to improve its products and services.

IP-verifiable
meta-externalagentTraining
Meta

Meta's crawler that collects public pages to index and improve its AI systems.

IP-verifiable
KimiBotTraining
Moonshot AI

Moonshot AI's bot for collecting public content that may train a later Kimi model.

IP-verifiable
BytespiderTraining
ByteDance

ByteDance's crawler, widely reported to collect public content for model training.

UA only
ERNIEBotTraining
Baidu

Baidu runs this to collect public pages that may help train its ERNIE models.

UA only
QwenBotTraining
Alibaba

Alibaba's crawler for Qwen. It gathers public content that may become training data.

UA only
ChatGLM-SpiderTraining
Zhipu AI

Zhipu AI sends this to gather public content that may go into training ChatGLM.

UA only
DeepSeekBotTraining
DeepSeek

DeepSeek's crawler for pulling public web content that may feed a future model.

UA only
cohere-aiTraining
Cohere

One of Cohere's crawlers, gathering public content for its AI systems.

UA only
cohere-training-data-crawlerTraining
Cohere

Cohere's crawler named for its purpose: gathering public content that may become model training data.

UA only
AI2BotTraining
Ai2

The Allen Institute's crawler (Ai2), gathering public documents to feed its AI research systems.

UA only
CCBotTraining
Common Crawl

Common Crawl's bot. It archives public pages into a free, open dataset that many model builders later train on.

IP-verifiable
OAI-AdsBotOther
OpenAI

OpenAI's crawler for ad and landing-page fetches, kept separate from its search and answer bots.

UA only
GrokBotOther
xAI

One of xAI's crawlers. xAI has not published what GrokBot fetches, so treat it as unverified xAI traffic for now.

UA only
xAI-BotOther
xAI

An xAI crawler with no public documentation on what it does. Its role stays unexplained until xAI says more.

UA only
xAI-GrokOther
xAI

A Grok-branded fetcher from xAI, with no published spec. Its exact purpose is unconfirmed.

UA only
xAI-Web-CrawlerOther
xAI

xAI's generically named web crawler. xAI hasn't said what it collects, so consider it undocumented traffic.

UA only
GrokOther
xAI

A plain Grok user agent from xAI, undocumented and easy to spoof, so treat it as unverified until it proves itself.

UA only
meta-externaladsOther
Meta

Meta's crawler for advertising and business products, not for answers or search.

IP-verifiable
facebookexternalhitOther
Meta

The Meta fetcher that builds link previews when someone shares your URL on Facebook, Instagram, or Messenger.

IP-verifiable
FacebookBotOther
Meta

A Meta crawler that turns up outside the usual link-preview flow. Meta files it under its general Facebook fetches.

IP-verifiable
DoubaobotOther
ByteDance

ByteDance's crawler tied to Doubao, its consumer AI assistant. ByteDance has published little about what it fetches.

UA only
YiyanBotOther
Baidu

Baidu's crawler linked to Yiyan, its AI assistant. Baidu documents it thinly, so specifics are limited.

UA only
TongyiBotOther
Alibaba

Alibaba's crawler for Tongyi, its AI assistant family. Alibaba publishes little detail about how it behaves.

UA only
AliyunBotOther
Alibaba

A crawler from Alibaba Cloud (Aliyun), with little public documentation about its exact purpose.

UA only

A note on freshness: IP ranges shown here are samples from the providers’ published lists as of July 2026. The linked official endpoints are the source of truth; ranges change without notice, and we would rather point you at the live list than pretend a snapshot is one.

Knowing them is half the job

The other half is seeing them on your own site. Voris AI visibility tells humans from AI agents in your traffic, verified, on every plan. Cookieless, with a free plan and a 14-day trial that needs no card.