CrawlPact

Insight · Publisher policy

Cloudflare Content Signals use=reference: What immediate, reference and full Mean

Cloudflare's Content Signals use field caps reuse at immediate, reference or full, separate from search, ai-input and ai-train. What it means, and its limits.

By CrawlPactPublished
Diagram: Content Signals has two separate dimensions. Purpose (search, ai-input, ai-train) runs across; the use ceiling (immediate, reference, full) runs up. use=reference is highlighted as retain, index and cite. Purpose and use are separate policy dimensions.

Content Signals began with three purpose signals in robots.txt: search, ai-input and ai-train, each yes or no. On 1 July 2026 Cloudflare started testing a fourth, optional field, use, which describes something different: how much a crawler may keep and reuse after it has accessed your content. It takes one of three values — immediate, reference or full — and Cloudflare now adds use=reference to the managed robots.txt it serves for customers who turned that feature on.

Key takeaway. Purpose and use are two separate dimensions. ai-train=no answers “for what purpose?”; use=reference answers “up to what level of reuse?”. Neither replaces the other, and neither is a technical block.

Where use stands today

Two facts set the frame. First, Cloudflare describes use as an extension it is testing: “we’re testing a new signal, use, that extends Content Signals”. Second, as of 2 October 2026 the Content Signals Policy site, contentsignals.org, still defines only the three purpose signals and does not document use. Treat use as a Cloudflare-proposed extension that one major operator has implemented, not as a settled part of the standard.

How to write the declaration itself, in robots.txt or elsewhere, is covered in how to set Content Signals. This article is about what the new field means.

Two dimensions, not one

Dimension Field Question it answers Values
Purpose search May content be used to build a search index? yes / no
Purpose ai-input May content be input to AI models at query time? yes / no
Purpose ai-train May content be used to train or fine-tune models? yes / no
Use ceiling (testing) use Up to what level may content be kept and reused? immediate / reference / full

Cloudflare’s managed line shows the two working together:

User-Agent: *
Content-signal: search=yes, ai-train=no, use=reference
Allow: /

Read literally, that says: search use is allowed, training is not, and whatever is allowed should stay at or below reference-level reuse. ai-input is not set, which under Cloudflare’s policy text “neither grants nor restricts permission”.

The three levels — in two wordings

Cloudflare describes the levels in two places, and the wording is not identical:

Level Announcement and robots.txt docs Browser Run /crawl docs
immediate “interact, but store and reuse nothing” “Ephemeral, single-response use, where content is not retained.”
reference “index, excerpt, and link back” “Content may be retained, indexed, or cited.”
full “summarize and reproduce” “Unrestricted use, including AI training.”

The two agree on the order — immediate is the narrowest, full the broadest — and broadly on reference: keep enough to index, quote briefly and link or cite. They differ on full: one wording is about summarising and reproducing, the other explicitly includes AI training. That difference is why the purpose signal still matters. A publisher that does not want training should say ai-train=no, not rely on a use level to imply it.

reference is the middle ground many publishers want: discoverable, linkable and citable, without granting summarisation, reproduction or unrestricted reuse. immediate goes further — interact, keep nothing — and full is the least restrictive ceiling.

One crawler that enforces it

Cloudflare’s own Browser Run /crawl endpoint shows how a crawler can honour the field. Since 31 August 2026 a caller declares a contentUse of reference or full (default full), and the crawl is rejected with a 400 error when that declared level “is more permissive than the site’s declared use level”:

  • the site sets use=full, or no use at all — any contentUse is allowed;
  • the site sets use=reference — contentUse: "reference" is allowed, the default full is not;
  • the site sets use=immediate — every crawl is rejected, because /crawl stores what it fetches, which Cloudflare says “is not compatible with ephemeral, single-response use”.

The same endpoint enforces the purpose signals: by default it declares all three purposes, so a site’s ai-train=no rejects the crawl unless the caller narrows its declared purposes to, say, search. And Cloudflare is explicit that this is a matter of trust: “Content Signals are trust-based. By setting crawlPurposes and contentUse, you are declaring to the site owner how you intend to use the crawled content.”

Cloudflare has also said it is building Bot Management rules that select and block bots by content use, is tracking content use for bots in its BotBase directory, and that “bots that reproduce in full cannot have the Verified status”. Those are Cloudflare’s own programmes; they do not make other crawlers honour the field.

Declaration, compliance, enforcement and rights are four layers

Layer Question What a use=reference line does
Declaration What does the site say it permits? States the ceiling publicly, in machine-readable form
Compliance Does a given crawler read and honour it? Depends on the crawler; Cloudflare’s /crawl documents that it does
Enforcement Is a non-compliant request actually stopped? Nothing by itself — Cloudflare says robots.txt compliance is voluntary
Rights What is the legal effect? Cloudflare’s policy text frames restrictions as express reservations of rights under Article 4 of EU Directive 2019/790; how that applies to a given case is a legal question, not a technical one

Keeping these apart prevents both mistakes: treating the line as a lock, and treating it as meaningless because it is not one.

Managed robots.txt is changing

Cloudflare announced on 15 September 2026 that Managed robots.txt will be deprecated in favour of Bot Preference Sync, which writes robots.txt groups from your Search, Training and Agent settings. Cloudflare’s published Bot Preference Sync example shows User-agent and Disallow groups; its posts do not say whether the use field carries over. Check your live robots.txt after any migration rather than assuming the Content-signal line is still there.

What CrawlPact checks today

CrawlPact’s Content Signals checker reads a Content-Signal HTTP response header on a site’s homepage and recognises search, ai-input and ai-train. Any other field in that header — including use — is preserved and shown as an unrecognised field, not interpreted. It does not currently read Content-signal lines inside robots.txt, which is where Cloudflare and contentsignals.org document them; for that form, read your live robots.txt directly. CrawlPact does not turn a trust-based signal into enforcement, decide its legal effect, or confirm which crawlers honour it.

The takeaway

use=reference is a ceiling on reuse, not a replacement for search, ai-input or ai-train. It gives publishers a way to say “you may index, excerpt and cite this — not more” in the same line that already states purposes. It is still a Cloudflare test, with one documented enforcing crawler. Set purposes explicitly, add a use ceiling if it matches your intent, verify the live robots.txt, and use enforcement controls when a preference is not enough — see Cloudflare’s Disallow AI Training setting for how Cloudflare pairs preference with blocking.

Sources

Claims about vendors, protocols and standards trace to these sources, last verified . Statements about CrawlPact's own systems are CrawlPact's first-party observations.

Check what your website currently declares to crawlers

CrawlPact audits the public crawler-policy signals a site exposes — robots.txt groups, robots meta and X-Robots-Tag directives, and related declarations — against a source-backed crawler registry. It does not guarantee crawler compliance, search ranking or AI citation, and it does not measure actual crawler traffic.

Audit a domain

Spotted an error or an outdated source? Report it through CrawlPact's corrections process.

Analytics preferences

CrawlPact uses optional Google Analytics and Microsoft Clarity on public pages to learn which content is useful. Clarity records clicks and scrolling (session replay), with anything you type masked. Neither runs in the app or admin areas, and CrawlPact works the same whether you accept or decline. See the privacy policy.