CrawlPact

Policy Observatory research methodology

Current methodology version: policy-observatory-methodology-v1. A material methodological change creates a new version; this page always describes the current one. Published research always states which version it used.

Two different "Observatory" concepts

CrawlPact's product interface uses an "Evidence Observatory" visual language (Observed → Interpretation → Impact → Action → Evidence) to present one domain's audit results. ThePolicy Observatory described on this page is a different, separate concept: a public research and measurement layer about the crawler-policy ecosystem as a whole, not about any one domain's audit.

Research authority principle

CrawlPact earns research authority by making evidence, methodology, denominators, limitations, provenance, and corrections visible — not by using authoritative-sounding language. Publications avoid unsupported phrasing such as "the definitive AI crawler database," "most websites block AI crawlers," or "industry standard" unless the underlying study design genuinely supports the exact claim.

Registry source policy

Every Registry Observatory figure is computed from CrawlPact's immutable crawler registry releases, never from the live/mutable crawler table and never from the static crawler directory's Markdown content. See how registry entries are verified for the underlying source-verification process that populates the registry itself.

Claim classes

Every publication statement falls into one of:

Observed
Directly measured — e.g. "8 of 23 tracked crawler records are classified as Search."
Computed
Deterministically calculated from observed data — e.g. a percentage.
Supported interpretation
An interpretation consistent with the measured signals, stated as such.
Unsupported
Not published, by definition — see the correction policy if one ever slips through.

Missing and unavailable data

Missing data is never silently treated as "absent" or "false." Where a Registry Observatory figure has no value (for example, a release published before checksums existed), the publication states that explicitly rather than omitting or zero-filling it.

Small-cell suppression

Registry Observatory counts are exhaustive population counts of CrawlPact's own registry, not a sample — the small-cell threshold does not apply to them. A future website-policy study would suppress any cohort smaller than 20 observations, showing "Insufficient sample" instead of an unstable percentage.

Comparability and trends

Two release-history points are described as a single "change since previous release," not a "trend," until enough comparable published releases exist. Registry taxonomy changes (a new purpose value, a new lifecycle status) are called out explicitly rather than folded silently into a raw count.

Website Policy Observatory — not yet built

A website-level policy benchmark (how commonly studied sites address AI training crawlers, and similar questions) requires a separately governed, versioned, bounded research corpus of publicly accessible websites, its own crawling policy, a capacity measurement, and a privacy review — none of that exists yet. Rather than publish a weak or unrepresentative statistic, only the Registry Observatory is published this phase.

Reproducibility

Every published statistic is generated by code from a pinned registry release — never hand-typed. Each publication carries a SHA-256 checksum over its content, independently recomputable from the same pinned release at any later time.

Corrections

A published statistic is never silently rewritten. A correction records what changed, why, the date, and whether the publication's conclusions changed, and remains visible on the publication page. See how to report a correction.

Analytics preferences

CrawlPact uses optional Google Analytics on public marketing pages to understand which content is useful. It is not used in the authenticated app or admin areas. You can accept or decline analytics without affecting the service.