CrawlPact

AI crawler directory

Reference pages for AI crawlers CrawlPact's registry tracks, each with its purpose, official source, and last-verified date. This is a growing, versioned registry — see methodology for how entries are verified.

22 crawlers from 8 operators · most recently verified 9/8/2026

How entries are verified

Each entry requires a reliable official source, a verified user-agent token, a purpose classification, and a verification date before publication. Entries are re-checked against each operator's own current documentation on a rolling basis, not left to go stale indefinitely — see the last-verified date on each crawler's page. Where an official source could not be reliably fetched for a verified crawler, that is disclosed on its own registry record rather than published as though automatically confirmed.

Search

  • Amzn-SearchBot

    Improves search experiences in Amazon products and services; Amazon's own documentation states it is not used for generative AI model training.

  • Claude-SearchBot

    Navigates the web to improve the relevance and accuracy of Claude's search results.

  • Googlebot

    Google's primary web crawler for Search indexing — not an AI-training-specific crawler.

  • Meta-WebIndexer

    Navigates the web to improve Meta AI search result quality.

  • OAI-SearchBot

    Used to discover and surface links to websites in ChatGPT search results.

  • PerplexityBot

    Indexes web content to power Perplexity's AI-generated search answers; Perplexity's own documentation recommends allowing it in robots.txt and states it respects robots.txt rules.

Training

  • Applebot-Extended

    Controls use of website content for training Apple Intelligence and other Apple generative AI models.

  • ClaudeBot

    Used by Anthropic to crawl publicly accessible content for model training.

  • Google-Extended

    Controls use of website content for training Gemini and Vertex AI generative models, independent of Search indexing.

  • GPTBot

    Used by OpenAI to crawl publicly accessible web content that may be used to train future models.

  • Meta-ExternalAgent

    Used by Meta to crawl content for training AI models and improving AI products.

User-triggered

  • Amzn-User

    Fetches a page on behalf of an end user or Amazon application, such as responding to an Alexa query that needs up-to-date information; not used for generative AI model training; Amazon's own documentation states it may not follow all robots.txt directives.

  • ChatGPT-User

    Fetches a web page in direct response to a user's question inside ChatGPT or a Custom GPT.

  • Claude-User

    Fetches a web page when a person directs Claude to access it as part of a query.

  • Perplexity-User

    Fetches a page in direct response to a user's question inside Perplexity; Perplexity's own documentation states it generally ignores robots.txt rules.

Agent / action

  • Google-CloudVertexBot

    Crawls requested by site owners for building Vertex AI Agents.

  • Meta-ExternalFetcher

    Fetches individual links at a user's request to support agentic AI capabilities in Meta products.

Advertising / validation

  • Meta-ExternalAds

    Crawls the web for use cases such as improving advertising and other business-related products.

  • OAI-AdsBot

    Validates the safety and relevance of web pages submitted as ads on ChatGPT — not used for AI training.

Research

  • CCBot

    Builds the open Common Crawl web corpus, which is reused by many third-party model trainers.

Mixed

  • Amazonbot

    Used by Amazon to improve its products and services, including Alexa answers, and may be used to train Amazon AI models.

Unknown

  • GoogleOther

    A generic Google crawler various internal product teams may use — Google's own documentation does not specify which teams or purposes.

Analytics preferences

CrawlPact uses optional Google Analytics and Microsoft Clarity on public marketing pages to understand which content is useful and how visitors actually use it. Neither is used in the authenticated app or admin areas. You can accept or decline analytics without affecting the service.