robots.txt and noindex in Practice: A CrawlPact Bing Site Scan Case Study
By CrawlPact · Published · Sources verified

A crawler-control rule can be valid and still be wrong for a URL it matches. CrawlPact — a company whose product audits other sites’ crawler policies — learned that from a Bing Webmaster Tools Site Scan of its own site in September 2026.
The scan covered the 81 URLs in CrawlPact’s sitemap and reported 1 error and 97 warnings. The
error was the important one: the public /audit/ landing page — listed in the sitemap
and meant to be indexed — was blocked by robots.txt.
This is a record of what happened, why, and what changed. For the general rules on choosing
between robots.txt, the robots meta tag and X-Robots-Tag, CrawlPact’s
decision guide is the reference.
The root cause.
Disallow: /audit/had been added to keep crawlers away from individual audit-result URLs under/audit/<id>. Robots rules match by path prefix, so the same rule also matched the landing page/audit/itself.
The symptoms: signals that contradicted each other
Before the fix, the landing page sent every “normal public page” signal: HTTP 200, a
self-referencing canonical URL, a place in sitemap.xml, and no noindex. At the same time,
robots.txt told every compliant crawler not to request anything under /audit/. The sitemap
said “discover this”; the crawl policy said “don’t fetch it.” Bing’s Site Scan reported the
contradiction directly.
The result pages were a different class of URL. A result at /audit/<id> was never meant to
appear in search, and already carried noindex twice — as an X-Robots-Tag: noindex, nofollow, noarchive response header and as a robots meta tag in the page. Result URLs were never part of
the sitemap.
There was a second copy of the same mistake inside the application. A route classifier used
startsWith("/audit/") to decide which responses get the noindex header — and "/audit/" also
starts with "/audit/". In production the landing page escaped only because it is served as a
prerendered static file that doesn’t pass through that code; any request path that did would have
stamped the public landing page noindex. The root problem was the same in both places: nothing
distinguished the exact landing page from the family of result pages beneath it.
The misconception: robots.txt is not an indexing control
robots.txt controls whether a crawler may fetch a URL. noindex controls whether a fetched page
may appear in results. They are easy to blur because both “keep things out of search” in casual
speech, but they work at different stages.
Google’s documentation is blunt: robots.txt “is not a mechanism for keeping a web page out of
Google.” For noindex to work, “the page or resource must not be blocked by a robots.txt file”;
otherwise “the crawler will never see the noindex rule, and the page can still appear in search
results, for example if other pages link to it.” Bing’s Webmaster Guidelines draw the same line:
“robots.txt controls crawl access, not indexing,” and NOINDEX is what to use “when a URL should
NOT appear in Bing search, Copilot experiences, or grounding API results.”
So the original rule did not just block the landing page by accident. For the result pages it was
meant to protect, it also hid the noindex those pages already carried.
Implementation rule. If you want a search crawler to obey a
noindexon an HTML page, do not also block that crawler from fetching the page inrobots.txt. It has to read the page to see the instruction.
The correction: separate the two URL classes
The obvious patch — keep the block and add an exception such as Allow: /audit/$ for the landing
page — was considered and rejected. It would have fixed Bing’s error while leaving the result
pages uncrawlable, so their noindex would still never be seen. Instead, each URL class got
signals that match its role:
- The public
/audit/landing page stays a normal indexable page: 200, self-canonical, listed in the sitemap, crawlable. - The
Disallow: /audit/rule was removed fromrobots.txtentirely. - Result pages at
/audit/<id>keepnoindexin both the response header and the page metadata — and are now crawlable, so crawlers can actually read it. - Result pages stay out of the sitemap.
- The route classifier matches only real results — a path under
/audit/that is longer than/audit/itself and isn’t a static-file alias of the landing page. - Automated tests encode the relationship, so sitemap,
robots.txtandnoindexrules cannot drift apart silently again.
Steps 5 and 6 turned a one-off fix into an engineering improvement. The goal was not to make one scanner stop complaining; it was to make the intended relationship impossible to break by accident.
The regression guard
CrawlPact’s test suite now rebuilds the sitemap’s route set from the same sources the sitemap
generator reads and evaluates the real production robots.txt against every URL — with
CrawlPact’s own RFC 9309 parser, the one the product uses on customer sites — for the generic *
group, Googlebot and bingbot. A browser-level test then walks the served sitemap and requires every
listed URL to return 200 with no redirect, be self-canonical, carry no noindex, nocache,
noarchive or nosnippet in either header or meta tag, and be allowed for all three crawler
tokens.
The result-page family is tested separately: result URLs must never enter the sitemap and must
carry noindex. One generic test cannot prove that two different URL classes are each configured
correctly; two targeted ones can.
What happened after the fix
The remediation went to production on 29 September 2026. A sitemap-wide audit of the live site
then found zero failures across all 81 public URLs: every one returned 200, was
self-canonical, avoided noindex, and was allowed by robots.txt for *, Googlebot and bingbot.
On 30 September, Bing’s robots.txt tester reported /audit/ as Allowed, URL Inspection
reported that the URL “can be indexed by Bing,” and a fresh Site Scan of the sitemap returned:
| Bing Site Scan | Before (29 Sep 2026) | After (30 Sep 2026) |
|---|---|---|
| Pages scanned | 81 | 81 |
| Errors | 1 | 0 |
| Warnings | 97 | 81 |
| Blocked by robots.txt | 1 (/audit/) |
0 |
| Title too long | 17 | 0 |
| “Alt attribute missing” | 80 | 81 |
The title-length warnings were a separate finding fixed in the same release. The remaining
warnings are a known scanner mismatch, not missing alt text: every page carries a small
decorative logo marked alt="" beside the visible “CrawlPact” wordmark, which accessibility
guidance treats as correct and Bing’s heuristic reports as missing. The count rose from 80 to 81
because Bing could now crawl /audit/ as well. Google Search Console’s live test also reported the
page as available to Google and indexable.
The outcome was not “Bing prefers noindex to robots.txt.” It was that each URL class finally
sent signals consistent with its role.
Five lessons that generalise
1. Test path rules against the exact parent page
A prefix rule aimed at descendants also matches the parent. Before deploying Disallow: /private/,
test a representative child URL and the exact /private/ page. The crawler matches strings, not
intentions.
2. Treat the sitemap, robots.txt and robots meta as one system
A sitemap entry says “this is a canonical public page worth discovering.” A Disallow for the same
URL says “don’t fetch it.” They can technically coexist, but for an indexable page they almost
always mean an architectural mistake.
3. A noindex page generally has to be crawlable
If a page is blocked from crawling, the engine may never see its noindex, and the bare URL can
still be indexed from links. For ordinary HTML that should stay reachable but out of search, make
the instruction observable.
4. noindex is not access control
noindex is a search directive, not a permission system: anyone with the URL can still open the
page if the server allows it. Google’s own advice for keeping something out of search lists
password protection alongside noindex. Genuinely private or sensitive content needs real
authentication; search directives and security are separate layers.
5. Automate the contradiction check
Manual review caught this once. A test catches it on every build: compare the URLs the sitemap
calls public with the robots.txt and meta rules that govern them, and fail before release.
A simple model for public, noindex and private URLs
| URL type | robots.txt | Index directive | Sitemap |
|---|---|---|---|
| Public canonical page | Allow crawl | Indexable (no noindex) |
Include |
| Reachable but intentionally excluded from search | Usually allow crawl so noindex can be seen |
noindex |
Exclude |
| Genuinely private or sensitive | Access control is primary; robots is secondary | Search directives are not protection | Exclude |
| Deleted or moved | Per migration plan | Return 404/410 or a permanent redirect | Remove old URL |
Why this matters for AI search too
The distinction now reaches beyond classic results. Bing’s guidelines say NOINDEX keeps a URL out
of Copilot experiences and grounding results as well as Bing search. Google’s robots meta
specification says nosnippet “will also prevent the content from being used as a direct input for
AI Overviews and AI Mode.” Search directives govern several discovery surfaces at once, which makes
an accidental conflict more expensive.
The safe approach is explicit intent: decide which URLs are public and indexable, which are
reachable but excluded, and which are private — then make routing, sitemap, robots.txt and
page-level directives all say the same thing.
The CrawlPact lesson
The Site Scan did not expose an exotic AI-search problem. It exposed a basic systems problem: two valid mechanisms combined in a way that contradicted the purpose of one URL. That is why technical SEO benefits from deterministic checks rather than assumptions.
CrawlPact’s product makes the same distinction in its own domain. A crawler-policy declaration is only useful once you know which crawler, purpose and resource it applies to. Broad rules are convenient; precise rules are safer — including on the site that publishes them.
Sources
Claims about vendors, protocols and standards trace to these sources, last verified . Statements about CrawlPact's own systems are CrawlPact's first-party observations.
- Introduction to robots.txtGoogle Search Central · Official documentation
- Block Search indexing with noindexGoogle Search Central · Official documentation
- Robots meta tag, data-nosnippet, and X-Robots-Tag specificationsGoogle Search Central · Official documentation
- Bing Webmaster GuidelinesBing Webmaster Tools Help · Official documentation
- CrawlPact robots.txt (current, corrected)CrawlPact · CrawlPact first-party evidence
- CrawlPact sitemap.xml (current)CrawlPact · CrawlPact first-party evidence
Related CrawlPact resources
See exactly which robots.txt rule matches each crawler
The free robots.txt validator parses a domain's robots.txt with the same RFC 9309 matching logic CrawlPact's scanner uses and shows the matched rule and line number for each documented crawler token. It cannot confirm that any crawler honours those rules.
Validate a robots.txt fileSpotted an error or an outdated source? Report it through CrawlPact's corrections process.