AI crawler policy report
sample-domain.example
Complete with warnings- 8/1/2026, 9:00:00 AM
- 2026.07.3
- 2026.07.2
- Allow search, block training
- Resource availability90
- Syntax & evaluation85
- Objective alignment45
- Cross-signal consistency70
Crawler access matrix
| Crawler | Operator | Purpose | Result | Matched rule | Line |
|---|---|---|---|---|---|
| OAI-SearchBot | OpenAI | search | Allowed | Allow: / | 4 |
| GPTBot | OpenAI | training | No explicit rule | — | — |
| ClaudeBot | Anthropic | training | Blocked | Disallow: / | 12 |
| Google-Extended | training | Blocked | Disallow: / | 12 | |
| ChatGPT-User | OpenAI | user triggered | Allowed | Allow: / | 4 |
| PerplexityBot | Perplexity | search | Mixed | Disallow: /private/ | 8 |
Findings
GPTBot has no explicit rule
robots.txt does not name GPTBot in any User-agent group, so no explicit rule currently applies to it.
Why it matters: The selected objective ("Allow search, block training") intends to restrict model-training crawlers, but an unaddressed crawler is not the same as an explicit block.
Recommended action: Add an explicit Disallow rule naming GPTBot's user-agent token.
CrawlPact cannot confirm whether GPTBot actually requests this website — only that no rule currently addresses it.
PerplexityBot is partially restricted
robots.txt disallows /private/ for PerplexityBot specifically, while allowing the rest of the site.
Why it matters: This may be intentional (protecting a private section) or an unintended narrower rule than other search crawlers receive.
Recommended action: Confirm whether restricting PerplexityBot from /private/ specifically, rather than all crawlers, is intended.
llms.txt
llms.txt /llms.txt
No /llms.txt file was found. This is informational — the file is optional.
llms-full.txt /llms-full.txt
No /llms-full.txt file was found. This is informational — the file is optional.
RSL (Really Simple Licensing)
No RSL declaration was found at /.well-known/rsl.xml. This is informational — RSL is optional and emerging.
RSL is a machine-readable declaration, not technical enforcement — see limitations.
Content Signals
No Content-Signal response header was present on the homepage. This is informational — Content Signals is an emerging, optional convention.
Related directives
- Meta robots tag: not present
- X-Robots-Tag header: not present
- Canonical URL: not present
Limitations
- This is a sample report generated from a fixture, not a real scan of any website.
- Results describe declared public policy signals; they do not prove actual crawler behaviour or compliance.
- CrawlPact does not have server-log access and cannot confirm whether any crawler actually requested this site.
- Crawler classifications reflect the registry version shown above and may change as operators update their documentation.
Ready to audit your own website?
Run a free public audit — no installation or server-log access required.
Audit a domain