Skip to content

Cloudflare’s Crawler Now Respects Content Signals

Cloudflare’s Browser Run crawler now respects the Content Signals use directive in robots.txt, giving web developers more control over how crawled content may be reused.

Cloudflare’s Crawler Now Respects Content Signals

On this page

Cloudflare has added a small but meaningful control to its Browser Run crawler: websites can now tell the crawler not only whether content may be accessed, but how much of that content may be reused. The change, released on August 31, 2026, makes Cloudflare’s crawler the first major implementation to enforce the newer use directive in Content Signals, giving web developers another way to separate ordinary crawling from broader content reuse.

The change adds a new layer to robots.txt

The familiar robots.txt file has traditionally been about access: a site can tell crawlers which paths they should or should not visit. Content Signals adds another question by describing the intended use of material after it has been collected. Its existing signals cover search, ai-input for supplying content to artificial intelligence systems at query time, and ai-train for training or fine-tuning models.

The new use signal goes a step further. It describes how extensively a crawler may retain and reuse content after accessing it. The current levels are reference and full, with reference being the more restrictive option. A publisher can therefore distinguish between allowing a crawler to index and reference material and allowing broader reuse of that material.

Cloudflare now checks that preference before crawling

Cloudflare’s Browser Run /crawl endpoint is designed to crawl a starting page and follow links across a site, returning content as HTML, Markdown, or JSON. The new implementation adds a contentUse parameter that tells the crawler what level of content use it intends to make. If that requested level is more permissive than the website allows in its robots.txt, Cloudflare says the crawl request is rejected with a 400 error rather than proceeding.

There is an important detail here: the parameter has a default of full. That means developers using the crawler without changing the setting are effectively declaring the broadest available use level. A crawler operator that only needs to collect material for reference can instead request reference, which may allow access to a site that does not permit full reuse.

What a web developer can actually control

Consider a publisher that wants its pages to remain discoverable while limiting how automated systems reuse the material. Its robots.txt can express Content Signals that separate search from artificial intelligence input and training. The crawler then compares those publisher preferences with the purposes declared by the crawler operator. This creates a more detailed conversation than simply blocking or allowing a bot.

For example, a site might allow search indexing but disallow artificial intelligence training. Cloudflare's crawler documentation already supports declaring individual crawl purposes, including search, ai-input, and ai-train. The new contentUse control deals with a different question: how much reuse is permitted after the crawler has obtained the content. Together, the controls give developers more precise ways to describe their publishing preferences.

The distinction matters for AI-powered web tools

This matters because modern web crawlers are no longer used only to build traditional search indexes. Cloudflare describes Browser Run as useful for applications such as research, monitoring, knowledge bases, and artificial intelligence systems that need current web information. A crawler can therefore retrieve a page for several very different reasons, and treating all of those activities as the same kind of access gives publishers little room to express what they actually want.

Content Signals was introduced by Cloudflare in 2025 as an extension to robots.txt intended to let publishers communicate those preferences. The system remains voluntary, however. The Content Signals project itself warns that crawler operators can technically ignore these instructions, and Cloudflare's documentation describes the signals as preference-based rather than a technical enforcement mechanism. That limitation is crucial: a signal can communicate a publisher's wishes, but it cannot magically stop every crawler from copying a page.

Cloudflare’s implementation is useful, but it is not universal enforcement

The practical value of the August 31 update is therefore narrower than a headline about stopping AI crawlers might suggest. It makes one widely used crawling service behave differently when a website publishes a restrictive use signal. It does not force unrelated crawlers, scrapers, or model operators to follow the same rules. Developers should still use conventional access controls when they need technical protection rather than a machine-readable preference.

There is also a subtle benefit for crawler developers. A crawler does not necessarily need full reuse permission for every job. A system collecting pages solely to create references may be able to request the less permissive reference level instead of declaring full. That makes the crawler's stated purpose more closely match the work it is actually performing, which is the kind of distinction the existing robots.txt model has struggled to express.

What developers should check now

Website owners using automated crawling should review their current robots.txt configuration rather than assuming that an existing rule covers every new form of content use. Cloudflare's documentation says its managed robots.txt feature can include Content Signals, and it has also introduced a content-use extension with levels such as immediate, reference, and full. Developers should decide whether those distinctions match the rights and business rules they want to communicate.

For developers building crawlers, the lesson is equally direct: declare the narrowest purpose and reuse level that genuinely matches the job. Cloudflare's new behavior shows that crawler APIs are beginning to treat these signals as part of the crawl request itself rather than as information that can simply be ignored. The bigger test will be whether other crawler operators adopt the same approach, because the value of a shared web standard depends on both sides speaking it.

M

Written by

M. Rizwan Mirza

I’m M. Rizwan Mirza, a Full Stack Developer with over 12 years of experience in web development and software solutions. I work with modern web technologies and enjoy building practical, reliable, and user-friendly digital solutions. I’m also part of TechWare House, where I work on web development projects and technology solutions. One of my favorite websites is TheQuranic.com. Through WizTechnoz, I share my knowledge, experience, tutorials, and useful insights about technology.

44 posts published

All posts by this author

0 Comments

No comments yet. Be the first to share your thoughts.

Join the conversation

Log in or create a free account to leave a comment. You can edit or delete your own comments any time.