Policy Observatory research methodology
Current methodology version: policy-observatory-methodology-v1. A material methodological change creates a new version; this page always describes the current one. Published research always states which version it used.
Two different "Observatory" concepts
CrawlPact's product interface uses an "Evidence Observatory" visual language (Observed → Interpretation → Impact → Action → Evidence) to present one domain's audit results. ThePolicy Observatory described on this page is a different, separate concept: a public research and measurement layer about the crawler-policy ecosystem as a whole, not about any one domain's audit.
Research authority principle
CrawlPact earns research authority by making evidence, methodology, denominators, limitations, provenance, and corrections visible — not by using authoritative-sounding language. Publications avoid unsupported phrasing such as "the definitive AI crawler database," "most websites block AI crawlers," or "industry standard" unless the underlying study design genuinely supports the exact claim.
Registry source policy
Every Registry Observatory figure is computed from CrawlPact's immutable crawler registry releases, never from the live/mutable crawler table and never from the static crawler directory's Markdown content. See how registry entries are verified for the underlying source-verification process that populates the registry itself.
Claim classes
Every publication statement falls into one of:
- Observed
- Directly measured — e.g. "8 of 23 tracked crawler records are classified as Search."
- Computed
- Deterministically calculated from observed data — e.g. a percentage.
- Supported interpretation
- An interpretation consistent with the measured signals, stated as such.
- Unsupported
- Not published, by definition — see the correction policy if one ever slips through.
Missing and unavailable data
Missing data is never silently treated as "absent" or "false." Where a Registry Observatory figure has no value (for example, a release published before checksums existed), the publication states that explicitly rather than omitting or zero-filling it.
Small-cell suppression
Registry Observatory counts are exhaustive population counts of CrawlPact's own registry, not a sample — the small-cell threshold does not apply to them. A future website-policy study would suppress any cohort smaller than 20 observations, showing "Insufficient sample" instead of an unstable percentage.
Comparability and trends
Two release-history points are described as a single "change since previous release," not a "trend," until enough comparable published releases exist. Registry taxonomy changes (a new purpose value, a new lifecycle status) are called out explicitly rather than folded silently into a raw count.
Website Policy Observatory — not yet built
A website-level policy benchmark (how commonly studied sites address AI training crawlers, and similar questions) requires a separately governed, versioned, bounded research corpus of publicly accessible websites, its own crawling policy, a capacity measurement, and a privacy review — none of that exists yet. Rather than publish a weak or unrepresentative statistic, only the Registry Observatory is published this phase.
Reproducibility
Every published statistic is generated by code from a pinned registry release — never hand-typed. Each publication carries a SHA-256 checksum over its content, independently recomputable from the same pinned release at any later time.
Corrections
A published statistic is never silently rewritten. A correction records what changed, why, the date, and whether the publication's conclusions changed, and remains visible on the publication page. See how to report a correction.