Insight · Publisher policy
Cloudflare Content Signals use=reference: What immediate, reference and full Mean
Cloudflare's Content Signals use field caps reuse at immediate, reference or full, separate from search, ai-input and ai-train. What it means, and its limits.

Content Signals began with three purpose signals in robots.txt: search, ai-input and
ai-train, each yes or no. On 1 July 2026 Cloudflare started testing a fourth, optional field,
use, which describes something different: how much a crawler may keep and reuse after it
has accessed your content. It takes one of three values — immediate, reference or full — and
Cloudflare now adds use=reference to the managed robots.txt it serves for customers who turned
that feature on.
Key takeaway. Purpose and use are two separate dimensions.
ai-train=noanswers “for what purpose?”;use=referenceanswers “up to what level of reuse?”. Neither replaces the other, and neither is a technical block.
Where use stands today
Two facts set the frame. First, Cloudflare describes use as an extension it is testing: “we’re
testing a new signal, use, that extends Content Signals”. Second, as of 2 October 2026 the
Content Signals Policy site, contentsignals.org, still defines only the three purpose signals and
does not document use. Treat use as a Cloudflare-proposed extension that one major operator has
implemented, not as a settled part of the standard.
How to write the declaration itself, in robots.txt or elsewhere, is covered in how to set
Content Signals. This article is about what the new
field means.
Two dimensions, not one
| Dimension | Field | Question it answers | Values |
|---|---|---|---|
| Purpose | search |
May content be used to build a search index? | yes / no |
| Purpose | ai-input |
May content be input to AI models at query time? | yes / no |
| Purpose | ai-train |
May content be used to train or fine-tune models? | yes / no |
| Use ceiling (testing) | use |
Up to what level may content be kept and reused? | immediate / reference / full |
Cloudflare’s managed line shows the two working together:
User-Agent: *
Content-signal: search=yes, ai-train=no, use=reference
Allow: /
Read literally, that says: search use is allowed, training is not, and whatever is allowed should
stay at or below reference-level reuse. ai-input is not set, which under Cloudflare’s policy text
“neither grants nor restricts permission”.
The three levels — in two wordings
Cloudflare describes the levels in two places, and the wording is not identical:
| Level | Announcement and robots.txt docs |
Browser Run /crawl docs |
|---|---|---|
immediate |
“interact, but store and reuse nothing” | “Ephemeral, single-response use, where content is not retained.” |
reference |
“index, excerpt, and link back” | “Content may be retained, indexed, or cited.” |
full |
“summarize and reproduce” | “Unrestricted use, including AI training.” |
The two agree on the order — immediate is the narrowest, full the broadest — and broadly on
reference: keep enough to index, quote briefly and link or cite. They differ on full: one
wording is about summarising and reproducing, the other explicitly includes AI training. That
difference is why the purpose signal still matters. A publisher that does not want training
should say ai-train=no, not rely on a use level to imply it.
reference is the middle ground many publishers want: discoverable, linkable and citable, without
granting summarisation, reproduction or unrestricted reuse. immediate goes further — interact,
keep nothing — and full is the least restrictive ceiling.
One crawler that enforces it
Cloudflare’s own Browser Run /crawl endpoint shows how a crawler can honour the field. Since 31
August 2026 a caller declares a contentUse of reference or full (default full), and the crawl
is rejected with a 400 error when that declared level “is more permissive than the site’s
declared use level”:
- the site sets
use=full, or nouseat all — anycontentUseis allowed; - the site sets
use=reference—contentUse: "reference"is allowed, the defaultfullis not; - the site sets
use=immediate— every crawl is rejected, because/crawlstores what it fetches, which Cloudflare says “is not compatible with ephemeral, single-response use”.
The same endpoint enforces the purpose signals: by default it declares all three purposes, so a
site’s ai-train=no rejects the crawl unless the caller narrows its declared purposes to, say,
search. And Cloudflare is explicit that this is a matter of trust: “Content Signals are
trust-based. By setting crawlPurposes and contentUse, you are declaring to the site owner how
you intend to use the crawled content.”
Cloudflare has also said it is building Bot Management rules that select and block bots by content use, is tracking content use for bots in its BotBase directory, and that “bots that reproduce in full cannot have the Verified status”. Those are Cloudflare’s own programmes; they do not make other crawlers honour the field.
Declaration, compliance, enforcement and rights are four layers
| Layer | Question | What a use=reference line does |
|---|---|---|
| Declaration | What does the site say it permits? | States the ceiling publicly, in machine-readable form |
| Compliance | Does a given crawler read and honour it? | Depends on the crawler; Cloudflare’s /crawl documents that it does |
| Enforcement | Is a non-compliant request actually stopped? | Nothing by itself — Cloudflare says robots.txt compliance is voluntary |
| Rights | What is the legal effect? | Cloudflare’s policy text frames restrictions as express reservations of rights under Article 4 of EU Directive 2019/790; how that applies to a given case is a legal question, not a technical one |
Keeping these apart prevents both mistakes: treating the line as a lock, and treating it as meaningless because it is not one.
Managed robots.txt is changing
Cloudflare announced on 15 September 2026 that Managed robots.txt will be deprecated in favour of
Bot Preference Sync, which writes robots.txt groups from your Search, Training and Agent
settings. Cloudflare’s published Bot Preference Sync example shows User-agent and Disallow
groups; its posts do not say whether the use field carries over. Check your live robots.txt
after any migration rather than assuming the Content-signal line is still there.
What CrawlPact checks today
CrawlPact’s Content Signals checker reads a Content-Signal
HTTP response header on a site’s homepage and recognises search, ai-input and ai-train.
Any other field in that header — including use — is preserved and shown as an unrecognised field,
not interpreted. It does not currently read Content-signal lines inside robots.txt, which
is where Cloudflare and contentsignals.org document them; for that form, read your live
robots.txt directly. CrawlPact does not turn a trust-based signal into enforcement, decide its
legal effect, or confirm which crawlers honour it.
The takeaway
use=reference is a ceiling on reuse, not a replacement for search, ai-input or ai-train. It
gives publishers a way to say “you may index, excerpt and cite this — not more” in the same line
that already states purposes. It is still a Cloudflare test, with one documented enforcing
crawler. Set purposes explicitly, add a use ceiling if it matches your intent, verify the live
robots.txt, and use enforcement controls when a preference is not enough — see Cloudflare’s
Disallow AI Training setting for how
Cloudflare pairs preference with blocking.
Sources
Claims about vendors, protocols and standards trace to these sources, last verified . Statements about CrawlPact's own systems are CrawlPact's first-party observations.
- Your site, your rules: new AI traffic options for all customersCloudflare Blog · Official documentation
- robots.txt setting (Content use signal)Cloudflare Docs · Official documentation
- /crawl endpoint: Content SignalsCloudflare Browser Run Docs · Official documentation
- Crawl endpoint now respects the Content Signals use directiveCloudflare Changelog · Official documentation
- Giving users choice with Cloudflare's new Content Signals PolicyCloudflare Blog · Official documentation
- Content Signalscontentsignals.org · Official documentation
- Have it both ways: stay discoverable in search while disallowing AI trainingCloudflare Blog · Official documentation
Related CrawlPact resources
- Cloudflare's Disallow AI Training Setting: Stay Discoverable in Search Without Allowing Training
- Web Bot Auth: Why User-Agent Strings Are Not Enough to Verify AI Agents
- How to set Content Signals: the robots.txt directive and the Content-Signal header
- RSL vs. Content Signals vs. robots.txt: which one wins when they disagree
- AI crawler policy on Cloudflare
Check what your website currently declares to crawlers
CrawlPact audits the public crawler-policy signals a site exposes — robots.txt groups, robots meta and X-Robots-Tag directives, and related declarations — against a source-backed crawler registry. It does not guarantee crawler compliance, search ranking or AI citation, and it does not measure actual crawler traffic.
Audit a domainSpotted an error or an outdated source? Report it through CrawlPact's corrections process.