CrawlPact

robots.txt and noindex in Practice: A CrawlPact Bing Site Scan Case Study

By CrawlPact · Published · Sources verified

Diagram: before the fix, Disallow: /audit/ blocked both the public /audit/ landing page listed in the sitemap and the /audit/<id> result pages, whose noindex was never seen. After removing the rule, /audit/ is crawlable, indexable and in the sitemap, and /audit/<id> is crawlable with noindex and never in the sitemap. robots.txt controls crawling; noindex controls indexing.

A crawler-control rule can be valid and still be wrong for a URL it matches. CrawlPact — a company whose product audits other sites’ crawler policies — learned that from a Bing Webmaster Tools Site Scan of its own site in September 2026.

The scan covered the 81 URLs in CrawlPact’s sitemap and reported 1 error and 97 warnings. The error was the important one: the public /audit/ landing page — listed in the sitemap and meant to be indexed — was blocked by robots.txt.

This is a record of what happened, why, and what changed. For the general rules on choosing between robots.txt, the robots meta tag and X-Robots-Tag, CrawlPact’s decision guide is the reference.

The root cause. Disallow: /audit/ had been added to keep crawlers away from individual audit-result URLs under /audit/<id>. Robots rules match by path prefix, so the same rule also matched the landing page /audit/ itself.

The symptoms: signals that contradicted each other

Before the fix, the landing page sent every “normal public page” signal: HTTP 200, a self-referencing canonical URL, a place in sitemap.xml, and no noindex. At the same time, robots.txt told every compliant crawler not to request anything under /audit/. The sitemap said “discover this”; the crawl policy said “don’t fetch it.” Bing’s Site Scan reported the contradiction directly.

The result pages were a different class of URL. A result at /audit/<id> was never meant to appear in search, and already carried noindex twice — as an X-Robots-Tag: noindex, nofollow, noarchive response header and as a robots meta tag in the page. Result URLs were never part of the sitemap.

There was a second copy of the same mistake inside the application. A route classifier used startsWith("/audit/") to decide which responses get the noindex header — and "/audit/" also starts with "/audit/". In production the landing page escaped only because it is served as a prerendered static file that doesn’t pass through that code; any request path that did would have stamped the public landing page noindex. The root problem was the same in both places: nothing distinguished the exact landing page from the family of result pages beneath it.

The misconception: robots.txt is not an indexing control

robots.txt controls whether a crawler may fetch a URL. noindex controls whether a fetched page may appear in results. They are easy to blur because both “keep things out of search” in casual speech, but they work at different stages.

Google’s documentation is blunt: robots.txt “is not a mechanism for keeping a web page out of Google.” For noindex to work, “the page or resource must not be blocked by a robots.txt file”; otherwise “the crawler will never see the noindex rule, and the page can still appear in search results, for example if other pages link to it.” Bing’s Webmaster Guidelines draw the same line: “robots.txt controls crawl access, not indexing,” and NOINDEX is what to use “when a URL should NOT appear in Bing search, Copilot experiences, or grounding API results.”

So the original rule did not just block the landing page by accident. For the result pages it was meant to protect, it also hid the noindex those pages already carried.

Implementation rule. If you want a search crawler to obey a noindex on an HTML page, do not also block that crawler from fetching the page in robots.txt. It has to read the page to see the instruction.

The correction: separate the two URL classes

The obvious patch — keep the block and add an exception such as Allow: /audit/$ for the landing page — was considered and rejected. It would have fixed Bing’s error while leaving the result pages uncrawlable, so their noindex would still never be seen. Instead, each URL class got signals that match its role:

  1. The public /audit/ landing page stays a normal indexable page: 200, self-canonical, listed in the sitemap, crawlable.
  2. The Disallow: /audit/ rule was removed from robots.txt entirely.
  3. Result pages at /audit/<id> keep noindex in both the response header and the page metadata — and are now crawlable, so crawlers can actually read it.
  4. Result pages stay out of the sitemap.
  5. The route classifier matches only real results — a path under /audit/ that is longer than /audit/ itself and isn’t a static-file alias of the landing page.
  6. Automated tests encode the relationship, so sitemap, robots.txt and noindex rules cannot drift apart silently again.

Steps 5 and 6 turned a one-off fix into an engineering improvement. The goal was not to make one scanner stop complaining; it was to make the intended relationship impossible to break by accident.

The regression guard

CrawlPact’s test suite now rebuilds the sitemap’s route set from the same sources the sitemap generator reads and evaluates the real production robots.txt against every URL — with CrawlPact’s own RFC 9309 parser, the one the product uses on customer sites — for the generic * group, Googlebot and bingbot. A browser-level test then walks the served sitemap and requires every listed URL to return 200 with no redirect, be self-canonical, carry no noindex, nocache, noarchive or nosnippet in either header or meta tag, and be allowed for all three crawler tokens.

The result-page family is tested separately: result URLs must never enter the sitemap and must carry noindex. One generic test cannot prove that two different URL classes are each configured correctly; two targeted ones can.

What happened after the fix

The remediation went to production on 29 September 2026. A sitemap-wide audit of the live site then found zero failures across all 81 public URLs: every one returned 200, was self-canonical, avoided noindex, and was allowed by robots.txt for *, Googlebot and bingbot.

On 30 September, Bing’s robots.txt tester reported /audit/ as Allowed, URL Inspection reported that the URL “can be indexed by Bing,” and a fresh Site Scan of the sitemap returned:

Bing Site Scan Before (29 Sep 2026) After (30 Sep 2026)
Pages scanned 81 81
Errors 1 0
Warnings 97 81
Blocked by robots.txt 1 (/audit/) 0
Title too long 17 0
“Alt attribute missing” 80 81

The title-length warnings were a separate finding fixed in the same release. The remaining warnings are a known scanner mismatch, not missing alt text: every page carries a small decorative logo marked alt="" beside the visible “CrawlPact” wordmark, which accessibility guidance treats as correct and Bing’s heuristic reports as missing. The count rose from 80 to 81 because Bing could now crawl /audit/ as well. Google Search Console’s live test also reported the page as available to Google and indexable.

The outcome was not “Bing prefers noindex to robots.txt.” It was that each URL class finally sent signals consistent with its role.

Five lessons that generalise

1. Test path rules against the exact parent page

A prefix rule aimed at descendants also matches the parent. Before deploying Disallow: /private/, test a representative child URL and the exact /private/ page. The crawler matches strings, not intentions.

2. Treat the sitemap, robots.txt and robots meta as one system

A sitemap entry says “this is a canonical public page worth discovering.” A Disallow for the same URL says “don’t fetch it.” They can technically coexist, but for an indexable page they almost always mean an architectural mistake.

3. A noindex page generally has to be crawlable

If a page is blocked from crawling, the engine may never see its noindex, and the bare URL can still be indexed from links. For ordinary HTML that should stay reachable but out of search, make the instruction observable.

4. noindex is not access control

noindex is a search directive, not a permission system: anyone with the URL can still open the page if the server allows it. Google’s own advice for keeping something out of search lists password protection alongside noindex. Genuinely private or sensitive content needs real authentication; search directives and security are separate layers.

5. Automate the contradiction check

Manual review caught this once. A test catches it on every build: compare the URLs the sitemap calls public with the robots.txt and meta rules that govern them, and fail before release.

A simple model for public, noindex and private URLs

URL type robots.txt Index directive Sitemap
Public canonical page Allow crawl Indexable (no noindex) Include
Reachable but intentionally excluded from search Usually allow crawl so noindex can be seen noindex Exclude
Genuinely private or sensitive Access control is primary; robots is secondary Search directives are not protection Exclude
Deleted or moved Per migration plan Return 404/410 or a permanent redirect Remove old URL

Why this matters for AI search too

The distinction now reaches beyond classic results. Bing’s guidelines say NOINDEX keeps a URL out of Copilot experiences and grounding results as well as Bing search. Google’s robots meta specification says nosnippet “will also prevent the content from being used as a direct input for AI Overviews and AI Mode.” Search directives govern several discovery surfaces at once, which makes an accidental conflict more expensive.

The safe approach is explicit intent: decide which URLs are public and indexable, which are reachable but excluded, and which are private — then make routing, sitemap, robots.txt and page-level directives all say the same thing.

The CrawlPact lesson

The Site Scan did not expose an exotic AI-search problem. It exposed a basic systems problem: two valid mechanisms combined in a way that contradicted the purpose of one URL. That is why technical SEO benefits from deterministic checks rather than assumptions.

CrawlPact’s product makes the same distinction in its own domain. A crawler-policy declaration is only useful once you know which crawler, purpose and resource it applies to. Broad rules are convenient; precise rules are safer — including on the site that publishes them.

Sources

Claims about vendors, protocols and standards trace to these sources, last verified . Statements about CrawlPact's own systems are CrawlPact's first-party observations.

See exactly which robots.txt rule matches each crawler

The free robots.txt validator parses a domain's robots.txt with the same RFC 9309 matching logic CrawlPact's scanner uses and shows the matched rule and line number for each documented crawler token. It cannot confirm that any crawler honours those rules.

Validate a robots.txt file

Spotted an error or an outdated source? Report it through CrawlPact's corrections process.

Analytics preferences

CrawlPact uses optional Google Analytics and Microsoft Clarity on public pages to learn which content is useful. Clarity records clicks and scrolling (session replay), with anything you type masked. Neither runs in the app or admin areas, and CrawlPact works the same whether you accept or decline. See the privacy policy.