CrawlPact

CrawlPact Research

AI Crawler Registry Landscape — September 2026

A source-backed snapshot of CrawlPact's verified AI-crawler registry as of release 2026.09.1: 24 tracked crawler records across 9 operators, classified by purpose, lifecycle status, and source-verification freshness.

Data collected
2026-09-20
Published
2026-09-20
Registry version
2026.09.1
Methodology
policy-observatory-methodology-v1

Key findings

Key findings with numerator, denominator, and scope
FindingCountScope
24 of 24 tracked crawler records are currently evaluation-eligible (active, deprecated, or replaced lifecycle status)24 / 24Registry release 2026.09.1
2 of 24 tracked crawler records are classified with purpose "advertising validation"2 / 24Registry release 2026.09.1
2 of 24 tracked crawler records are classified with purpose "agent"2 / 24Registry release 2026.09.1
2 of 24 tracked crawler records are classified with purpose "mixed"2 / 24Registry release 2026.09.1
1 of 24 tracked crawler records are classified with purpose "research"1 / 24Registry release 2026.09.1
7 of 24 tracked crawler records are classified with purpose "search"7 / 24Registry release 2026.09.1
5 of 24 tracked crawler records are classified with purpose "training"5 / 24Registry release 2026.09.1
1 of 24 tracked crawler records are classified with purpose "unknown"1 / 24Registry release 2026.09.1
4 of 24 tracked crawler records are classified with purpose "user triggered"4 / 24Registry release 2026.09.1
24 of 24 tracked crawler records carry a recorded source-verification date24 / 24Registry release 2026.09.1
0 of 24 evaluation-eligible crawler records are due for re-verification (never verified, or last verified more than 180 days ago)0 / 24Registry release 2026.09.1
9 distinct operators are represented across 24 tracked crawler records9 / 24Registry release 2026.09.1

Dataset

Every figure in this publication is computed directly from CrawlPact's immutable registry release "2026.09.1" (id 8d192f1d-3c44-4f83-9cea-32e56fc54092), published 2026-09-20T11:46:53.082Z. No customer domains, saved-domain portfolios, Agency client portfolios, anonymous audits, or product analytics data are used. This dataset is CrawlPact's own crawler-identity registry, not a sample of the web.

Methodology

Methodology version policy-observatory-methodology-v1. Counts are drawn from every crawler entry recorded in the pinned registry release's immutable snapshot (registry_version_entries), resolved via the same read path production evaluation and historical scan rendering use — never from the live, mutable crawlers table. See docs/research/RESEARCH_METHODOLOGY.md and docs/registry/PHASE_15_FULL_SOURCE_REVERIFICATION_REPORT.md for how each underlying crawler record was itself verified.

Results

2 registry releases have been published in total. The most recent published release ("2026.09.1") recorded 1 crawler addition(s), 0 removal(s), 0 purpose reclassification(s), 0 token-identity change(s), and 0 evidence-only source correction(s) (no evaluation impact) relative to the previous release. Fewer than 3 comparable releases exist, so this is reported as a single change-since-previous-release observation, not a trend.

Registry/ruleset provenance

Registry version: 2026.09.1 (checksum: 9a4f14303d99f24d8ba513d2d158263dd48df3710427fa3ac00065607bb9ed2b). This publication does not evaluate customer policy data, so no ruleset version applies.

Limitations

  • This publication describes CrawlPact's own crawler-identity registry, not a survey or sample of live websites — it does not measure how many websites block or allow any crawler.
  • Purpose classification reflects each operator's own published documentation as independently reverified by CrawlPact; it does not measure or predict actual crawler behaviour.
  • Registry release counts reflect what CrawlPact has verified and published — an operator may run undocumented or unverified crawlers not yet reflected here.
  • Verification-freshness figures describe when CrawlPact last reverified a source, not whether the underlying operator documentation has since changed without CrawlPact's knowledge.

Corrections

No corrections have been made.

MethodologyRegistry ObservatoryReport a correctionAudit a domain free

Analytics preferences

CrawlPact uses optional Google Analytics and Microsoft Clarity on public marketing pages to understand which content is useful and how visitors actually use it. Neither is used in the authenticated app or admin areas. You can accept or decline analytics without affecting the service.