ai-crawlers
web-infrastructure
robots-txt
publishing
jargon-dictionary

The Internet's Politest File Just Got a Doorman

robots.txt was always a request, never a rule. Cloudflare's Bot Preference Sync publishes your bot settings into the file itself, names four conditions a search-and-training crawler must meet to keep reading sites that refuse training, and flips the default for ad-funded newcomers.

Since 1994, the web's oldest defense against unwanted robots has been a text file politely asking them to leave. On 21 August, Cloudflare started writing that file for you. And on 15 September, for new ad-funded websites behind Cloudflare, the default answer to "may I train on this?" flips to no.

Prefer it read to you? Zara does voices now (12 min).

Zara is a character and this voice is synthesized. Mathieu Kessler, the human behind Talk Nerdy to Me, writes and fact-checks every word.

The note taped to every door on the internet

On the front of nearly every website, at the same fixed address, there is a note taped to the door. It has been there since 1994. You can go and read one right now: take any site you like and add /robots.txt to the end.

The note says which robots are welcome and which are not. That is all it does. It is a plain text file. It has no lock, no alarm, and no way of knowing whether anybody read it. A crawler that ignores it suffers no consequence from the file itself, because a request is not a rule.

For three decades this mostly worked, and it worked for a reason worth being precise about: the robots doing the reading were search engines, and a search engine wants to be let back in tomorrow. The note held because the visitors had manners, and the visitors had manners because they had something to lose.

Then a new kind of visitor arrived, one that only needs to read your site once.

Two things happened this summer that change what the note means, both of them from Cloudflare, the company whose network sits in front of a very large slice of the web. The second one landed on 21 August. I read both announcements so you do not have to.

Nerd to English

Four words you need, and then we can go.

A crawler is a program that visits web pages and reads them, automatically, at scale. Not evil, not good. Search engines run them, researchers run them, and AI companies run them. The question is never whether crawlers exist. It is what happens to your words after one has read them.

robots.txt is the note on the door. A text file at a fixed address that lists which crawlers may read which parts of a site. It is voluntary. It became an official internet standard only in 2022, twenty-eight years after everyone started using it, which tells you how much of the web runs on habit.

A user agent is the name a visitor announces at the door. Every crawler introduces itself with one. The whole polite system depends on visitors telling the truth about who they are, and on one name meaning one thing.

Training in this story means a crawler taking your content to teach an AI model. Cloudflare's taxonomy separates it from Search, which indexes your pages so people can find them, and from Agent, which is software fetching a page right now because a human asked. Same reading, three very different afterlives for your words.

What shipped on 21 August

On 21 August, Cloudflare shipped a feature called Bot Preference Sync, available on all Cloudflare plans. The mechanics are almost boringly simple: whatever you set in your Cloudflare dashboard about AI crawlers now gets written into your robots.txt for you, wrapped between two marker lines, prepended to whatever you already had, and refreshed as the list of known bots changes. Their stated goal fits in one sentence:

"[T]he preference you set is the preference you publish."
Cloudflare Blog, "Bot Preference Sync", 21 August 2026

Why would that need to exist? Because of a loophole so human it hurts. Lots of site owners had told Cloudflare's firewall one thing and their robots.txt another, usually because the file was set up years ago and nobody remembered it was there. And a mismatch, it turns out, is useful to exactly one party:

"When your stated preferences and your enforced rules disagree, some crawlers treat it as a basis to disregard your preferences."
Cloudflare Blog, "Bot Preference Sync", 21 August 2026

Read that the way I did. Some crawlers were treating a stale note as permission. The fix is not a better lock. The fix is making sure the note always says what the lock is doing, so nobody can claim they were confused.

One name at the door, two jobs inside

Now the real story, which is not the file. It is the bundling.

The most important crawlers on the internet do two jobs with one name. The same user agent that indexes your site for search, which you want, may also collect your pages for AI training, which you may not. One name, one door, two afterlives for your words.

That bundling had a consequence: almost nobody dared say no. Cloudflare's own July announcement is blunt about what happens if you do, naming names: customers who select block-Training end up blocking Googlebot, Applebot and BingBot. Blocking the training meant blocking the search traffic, and the search traffic pays the rent. So the choice on offer was never really training yes or training no. It was everything or nothing, and everything won by default.

The 21 August change makes the bundle negotiable. A crawler that does both search and training can keep reading sites that refuse training, but only by meeting four conditions:

It must respect a no-training preference in robots.txt, by any mechanism. It must give site owners a way to opt out of AI summaries. It must provide URL-level visibility into which pages were made available for training, alongside search metrics. And it must show publicly that disallowing training does not hurt your traditional search results.

That last one deserves a slow second read. It asks the crawler operator to prove the thing every publisher has been afraid of, which is that saying no to the training half quietly costs you the search half. And the enforcement for the whole list is flat:

"Crawlers that don't provide Transparency will not get the benefit of the doubt — they're still blocked when you disallow training."
Cloudflare Blog, "Bot Preference Sync", 21 August 2026

Transparency, in Cloudflare's phrase, is the price of admission. Who is paying it and who is not gets tracked publicly, on Cloudflare Radar, which is a report card for robots, published where everyone can see it.

15 September, the day the default flips

The second date to know is 15 September, and it comes from the July announcement.

From that day, new domains joining Cloudflare get new defaults: on pages that display ads, Training and Agent crawling are blocked by default. Search stays allowed. And a new customer who ticks the box saying they make money from ads on their domain gets Training set to Disallow before they have configured anything at all.

The reasoning is the most quotable thing in either post:

"An ad is a signal that a website owner meant for a person to land there and see it. [...] [O]n those pages, we treat human attention as the end goal."
Cloudflare Blog, "Content Independence Day", 1 July 2026

Sit with what that means at the scale of a network like this one. For three decades, the web's default answer to a visiting crawler has been yes, because saying no required knowing the note existed, knowing the crawler's name, and editing a file by hand. For ad-funded sites arriving at Cloudflare after 15 September, the default answer to the training question becomes no. Nobody passed a law. A very large infrastructure company changed a form.

What this proves, and what it does not

This is real, and half of it is already live. Both announcements are Cloudflare's own, with dates, mechanics and marker syntax. Bot Preference Sync shipped on 21 August; the September defaults are a dated, scoped change that has not happened yet. None of this is a proposal.

It only works at Cloudflare's door. robots.txt is still just a request, everywhere, and the file itself gained no new enforcement powers. What changed is that one doorman with a very long guest list now enforces what the note says, for the sites standing behind it. If your site lives elsewhere, nothing about your week changed.

The rules are one company's rules. The four conditions are written, judged and enforced by a single private company, which is now grading the internet's crawlers and publishing the grades. Whatever you think of the rules, nobody elected the doorman, and the doorman does not have to stand for re-election.

And the defaults are for newcomers. The 15 September flip is stated for new domains onboarding to Cloudflare, and Cloudflare says it will keep notifying existing customers before the deadline, with an opt-out for anyone who wants nothing to move. Your new neighbour's settings start out different from yours, and that is how defaults change a whole neighbourhood: one move-in at a time.

Why you, a person with a normal life, might care

Because you read the answers built from all this. Most of what an AI answer knows was assembled from pages someone wrote. The question of who may read those pages, for what purpose, on what terms, is being settled right now, not in a courtroom but in configuration defaults. The defaults are where the outcome actually lives, because almost nobody changes them.

Because if you have a site behind Cloudflare, your note may soon say things you did not type. Go and look at your robots.txt. If Bot Preference Sync is on, there will be a block between two BEGIN and END marker lines that your dashboard wrote on your behalf. It is your preference, published. Worth reading once, the way anything signed in your name is.

Because at this one very large door, the web is quietly splitting in two. Pages meant for human eyes, marked by ads, defaulting to no-training. And everything else. That split used to be a metaphor. As of this summer it is a checkbox with a default, which is the strongest form a metaphor can take.

The human part

The robots.txt convention was worked out in 1994, on a mailing list, by people who trusted each other. It is one of the oldest continuously working agreements on the internet, and for twenty-eight of its thirty-two years it was not even an official standard. It was just what everyone did.

It survived every wave of growth the web threw at it, because every visitor with scale had a reason to behave. What it could not survive unchanged was a visitor who benefits from reading once and never coming back.

So the politest file on the internet finally got a doorman. Not because manners stopped mattering, but because manners were the only mechanism, and somebody built a second one.

The note on the door did not change. Someone is finally standing next to it.

So: go and read a robots.txt today. Yours, your employer's, your favourite newspaper's. Add /robots.txt to the address and see who is welcome and who is named and turned away. If you find something strange in one, a bot named after a dead product, a plea written to no one, a joke a developer left in 2011, send it my way. Those I collect.

Sources

More Where This Came From

Plain-language translations of the machinery and the money behind the tech headlines. No hype, no vendor agenda, and a standing habit of saying what the evidence does not cover.