Audit and monitor your website's AI crawler policy.
CrawlPact independently audits robots.txt and related public signals, separates search from training and agent crawlers, explains conflicts, and monitors changes across any hosting provider.
No installation, server-log access, or AI API required. Try example.com.
- Vendor-neutral
- Evidence-based findings
- Works with any hosting provider
- Deterministic recommendations
One crawler-policy mistake can create the wrong outcome.
Search crawlers may be blocked unintentionally
A broad AI-crawler rule may also affect crawlers used for search or user-requested retrieval — not only model-training crawlers.
Training access may remain unspecified
A website may intend to restrict model-training crawlers while publishing no explicit rule for several verified crawler tokens. Unspecified is a neutral state, not a guarantee either way.
Deployments and CDN settings can change the public policy
The policy a team intends may differ from the public response after a deployment, framework update, or CDN-generated rewrite.
Not every website needs the same policy, and not all AI crawling is harmful — the point is knowing, with evidence, what your website currently declares.
See what a report looks like
Illustrative example using a synthetic demonstration domain — not a real scan result.
sample-domain.example (demonstration data)
AI crawler policy report
Scanned 01/08/2026 · Registry v2026.07.3
- Resource availability90
- Syntax & evaluation85
- Objective alignment45
- Cross-signal consistency70
OAI-SearchBot
OpenAI · Search
GPTBot
OpenAI · Training
ClaudeBot
Anthropic · Training
Google-Extended
Google · Training
Example finding
GPTBot has no explicit rule
robots.txt, fetched 2026-08-01 — no User-agent group names GPTBot.
Recommended action: Add an explicit Disallow rule naming GPTBot's user-agent token.
How CrawlPact works
- 1
Audit public policy signals
CrawlPact requests publicly accessible resources such as robots.txt and related policy signals.
- 2
Evaluate by crawler purpose
Verified crawler tokens are grouped by search, training, user-triggered retrieval, or agent use.
- 3
Explain findings with evidence
The report identifies explicit rules, unspecified policies, conflicts, unavailable resources, and recommended next steps.
- 4
Monitor changes
Saved domains can be compared against future website responses and crawler-registry updates on paid plans.
Public policy auditing does not require server-log access or installation. Actual crawler traffic is outside the audit's proof boundary — see methodologyand crawler directory.
Not every AI crawler serves the same purpose.
Crawler operators increasingly separate purposes into distinct tokens. Blocking one purpose does not automatically affect another, and classifications are reviewed against official documentation rather than assumed.
Search
Used to support search or discovery experiences, indexing content to help answer queries.
Training
Used to collect content for model development or training. A separate decision from search, often using a different crawler token.
User-triggered retrieval
Used when a person explicitly asks a product to retrieve a specific page in real time.
Agents
Used by automated or semi-automated agent workflows performing a task, distinct from a one-shot retrieval.
Registry classifications may change when an operator updates its own documentation — see the AI crawler directory for every verified entry and its source.
Built for how your team actually works
The same audit and monitoring engine, with guidance shaped around what your role actually needs to know.
Agencies
Agencies must understand, explain, and monitor AI crawler policy state across many client websites at once — one-off manual checks don't scale, and clients ask about it in language that doesn't map cleanly to a robots.txt file.
Publishers
Search and AI-training crawlers often warrant different decisions, but a site's public signals frequently don't distinguish them clearly — and CDN or deployment changes can alter the actual response without an editorial decision being made.
SaaS and documentation teams
Documentation discoverability decisions, the search-versus-training distinction, and deployment-platform-driven policy drift all intersect on the domains SaaS teams run — often across multiple subdomains with different owners.
Web developers
The crawler policy a developer intends to ship and the crawler policy a framework, hosting platform, or CDN actually serves can diverge — and that gap is invisible until someone checks the live, deployed response.
AI crawler directory
View all crawlersGovern crawler policy across every client website.
CrawlPact is built primarily for agencies and multi-site teams — save every client domain in one place, apply objective-based recommendations, and monitor for website or registry changes without exposing internal tooling to clients.
- 1Audit client domains
- 2Establish policy baselines
- 3Identify conflicts and unspecified rules
- 4Apply client-approved recommendations
- 5Monitor website and registry changes
- 6Share or export evidence
Portfolio monitoring
Group domains, batch-import a client list, and schedule recurring rechecks — plan limits apply, see pricing.
Client-ready sharing
Share a report with your own agency name and logo attached — CrawlPact's methodology and limitations always remain visible on a shared report, and branding never replaces them.
Independent evidence, not another hosting-provider control panel.
CrawlPact works across publicly accessible hosting setups, does not require a website to use one CDN, preserves the evidence behind every finding, uses a source-backed crawler registry, distinguishes website-policy changes from registry-driven changes, and publishes its own methodology and limitations.
CrawlPact audits the public policy signals a website publishes. It does not control external crawlers or guarantee that they will comply.
Supported public signals
Primary signals (robots.txt, HTTP headers, HTML meta directives) are checked on every scan. Optional or emerging signals are checked when publicly available and evaluated according to the current published methodology — none is required for AI visibility.
- robots.txt
- AI crawler user-agent groups
- llms.txt
- llms-full.txt
- RSL
- Content Signals
- Robots meta
- X-Robots-Tag
- Sitemap declarations
- Relevant HTTP headers
A correct policy today can change tomorrow.
Website-policy changes
- robots.txt, header, or meta directive changed
- A resource became unavailable
- A CDN rewrote the public response
- A deployment changed the effective policy
Registry-driven changes
- A new crawler token was added
- An operator changed its documented purpose
- Source verification changed
- Classification uncertainty changed
Monitoring frequency depends on your plan — see pricing. Changes surface through an in-app notification centre and a private Atom feed.
Pricing
View full pricingMonthly and yearly pricing. Billing is handled by Paddle; applicable taxes may be calculated during checkout.
Frequently asked questions
See what your website currently tells AI crawlers.
Run a free public-policy audit with no installation or server-log access.
Results describe published policy signals and do not guarantee crawler behaviour.
View a sample report