free reference
The AI crawler directory
Every AI assistant, search index, and training crawler that might visit a website: what it does, whether its traffic can be verified, and how to allow or block it. Built from the providers’ own published sources.
Look up any crawler
Search by crawler or provider, or filter by what the bot is for. Every entry links to its own page with verification guidance, sample IP ranges where the provider publishes them, and ready-made robots.txt lines.
All 62 crawlers
OpenAI's live fetcher. When someone asks ChatGPT a question your page can answer, this is what opens it in the moment.
IP-verifiableAnthropic's live fetcher for Claude. It reads your page the instant a person's question calls for it.
IP-verifiableThe fetch behind a Perplexity answer. It reads your page so Perplexity can quote it and list it as a source.
IP-verifiableOne of Google's user-triggered agents. It opens a page in the moment, acting for a person rather than fetching an answer.
IP-verifiableThe fetcher behind Google's NotebookLM. It opens a page when someone adds it as a source to research with.
IP-verifiableA Google fetcher for read-aloud and assistant features. It opens your page so the text can be spoken back.
IP-verifiableAnother label in Google's user-triggered agent family. Like Google-Agent, it fetches pages while carrying out a person's task.
IP-verifiableMistral's live fetcher for Le Chat. It opens your page when a person's question needs it.
IP-verifiableMicrosoft Copilot fetching live. It reads your page in the moment to answer whatever someone just asked.
IP-verifiableAmazon's live fetcher behind products like Alexa, opening your page to pull a fresh answer for someone.
IP-verifiableDuckDuckGo's real-time fetcher for DuckAssist. It reads your page so the AI answer can quote and cite it.
IP-verifiablexAI's live fetcher for Grok. It opens your page as Grok answers, though xAI publishes no IP ranges to confirm it.
UA onlyxAI's fetcher for Grok's deep search mode. It reads your page live while building a longer, researched answer.
UA onlyA Meta fetcher that opens a specific page when a person requests or shares its link inside a Meta product.
IP-verifiableMoonshot's live fetcher for Kimi. It opens your page when a question there needs an answer.
IP-verifiableAlibaba's live fetcher for Qwen, opening your page mid-answer. Alibaba publishes no IP ranges, so it's user-agent only.
UA onlyOpenAI's crawler that discovers and refreshes pages so they can turn up in ChatGPT search.
IP-verifiableAnthropic's crawler that indexes pages so Claude can find and cite them when it searches the web.
IP-verifiablePerplexity's indexing crawler, quietly keeping its answer engine's map of the web current.
IP-verifiableGoogle's inspection fetcher, the one behind Search Console's URL inspection and the Rich Results test.
IP-verifiableGoogle's main search crawler, indexing the web for Search. Cloudflare notes the same traffic can feed AI Overviews.
IP-verifiableMistral's indexing crawler, building the map Le Chat draws on when it answers.
IP-verifiableMicrosoft's long-running search crawler, indexing pages for Bing and, by extension, Copilot.
IP-verifiableA legacy Microsoft search crawler, largely superseded by Bingbot but still turning up in logs.
IP-verifiableAmazon's crawler that indexes pages to make them eligible for Amazon's search experiences.
IP-verifiableMeta's indexing crawler, working to sharpen the results Meta AI returns in search.
IP-verifiableMoonshot's search crawler for Kimi, indexing pages so the assistant can find them when answering.
IP-verifiableByteDance's crawler that indexes pages for TikTok's search and discovery surfaces.
UA onlyBaidu's main search crawler, indexing pages for China's largest search engine.
UA onlyYou.com's crawler, indexing pages for its search and answer engine.
UA onlyOpenAI's training crawler. It gathers public pages that may feed a future GPT model.
IP-verifiableAnthropic's crawler for gathering public web text that may go toward improving Claude.
IP-verifiableA catch-all Google crawler that product teams use to pull public pages for research and development.
IP-verifiableNot a crawler at all, but a robots.txt token that decides whether Google may use your content to train and ground Gemini.
robots.txt tokenGoogle's crawler for on-demand fetches a site owner requests while building a Vertex AI agent.
IP-verifiableA robots.txt token rather than a bot. It tells Apple whether it may use already-crawled content for AI training.
robots.txt tokenApple's crawler, which Cloudflare places across both search and training uses, feeding Siri, Spotlight, and Safari.
IP-verifiableAmazon's general crawler, gathering public content to improve its products and services.
IP-verifiableMeta's crawler that collects public pages to index and improve its AI systems.
IP-verifiableMoonshot AI's bot for collecting public content that may train a later Kimi model.
IP-verifiableByteDance's crawler, widely reported to collect public content for model training.
UA onlyBaidu runs this to collect public pages that may help train its ERNIE models.
UA onlyAlibaba's crawler for Qwen. It gathers public content that may become training data.
UA onlyZhipu AI sends this to gather public content that may go into training ChatGLM.
UA onlyDeepSeek's crawler for pulling public web content that may feed a future model.
UA onlyOne of Cohere's crawlers, gathering public content for its AI systems.
UA onlyCohere's crawler named for its purpose: gathering public content that may become model training data.
UA onlyThe Allen Institute's crawler (Ai2), gathering public documents to feed its AI research systems.
UA onlyCommon Crawl's bot. It archives public pages into a free, open dataset that many model builders later train on.
IP-verifiableOpenAI's crawler for ad and landing-page fetches, kept separate from its search and answer bots.
UA onlyOne of xAI's crawlers. xAI has not published what GrokBot fetches, so treat it as unverified xAI traffic for now.
UA onlyAn xAI crawler with no public documentation on what it does. Its role stays unexplained until xAI says more.
UA onlyA Grok-branded fetcher from xAI, with no published spec. Its exact purpose is unconfirmed.
UA onlyxAI's generically named web crawler. xAI hasn't said what it collects, so consider it undocumented traffic.
UA onlyA plain Grok user agent from xAI, undocumented and easy to spoof, so treat it as unverified until it proves itself.
UA onlyMeta's crawler for advertising and business products, not for answers or search.
IP-verifiableThe Meta fetcher that builds link previews when someone shares your URL on Facebook, Instagram, or Messenger.
IP-verifiableA Meta crawler that turns up outside the usual link-preview flow. Meta files it under its general Facebook fetches.
IP-verifiableByteDance's crawler tied to Doubao, its consumer AI assistant. ByteDance has published little about what it fetches.
UA onlyBaidu's crawler linked to Yiyan, its AI assistant. Baidu documents it thinly, so specifics are limited.
UA onlyAlibaba's crawler for Tongyi, its AI assistant family. Alibaba publishes little detail about how it behaves.
UA onlyA crawler from Alibaba Cloud (Aliyun), with little public documentation about its exact purpose.
UA onlyA note on freshness: IP ranges shown here are samples from the providers’ published lists as of July 2026. The linked official endpoints are the source of truth; ranges change without notice, and we would rather point you at the live list than pretend a snapshot is one.
Knowing them is half the job
The other half is seeing them on your own site. Voris AI visibility tells humans from AI agents in your traffic, verified, on every plan. Cookieless, with a free plan and a 14-day trial that needs no card.