Robots.txt files are getting rewritten on live sites right now, often by people who have rarely opened one before. For most of its thirty-plus years, it sat untouched. A developer set it up once and nobody thought about it again. That era is over. AI crawlers, answer engines, and user-prompted agents have turned a sleepy text file into the main policy document a site publishes about who gets to read it, for what, and under what terms.
The pressure is coming from both ends. Traffic patterns have shifted hard toward bots, and the bots themselves no longer fit the single category robots.txt was built to describe. A rewrite is happening whether site owners are driving it or not. For teams that want help thinking through which agents to allow, block, or rate-limit, working with an AI SEO agency that reads server logs weekly beats guessing from a template.
Phase One: The File Nobody Edited Became the File Everyone Argues About
Robots.txt started as a gentleman’s agreement. A webmaster listed which paths were off-limits, well-behaved crawlers obeyed, and misbehaved ones did whatever they wanted. The protocol was finally codified as RFC 9309 in 2022, which pinned down the syntax, caching behavior, and error handling while keeping the same voluntary spirit.
Nothing in the standard forces compliance. It defines what compliance looks like and leaves the rest to the crawler.
That worked when the crawlers were mostly search engines sending traffic back. The deal was simple: let us index you, we send you readers. When the crawlers started training models instead of sending readers, the deal fell apart. Site owners who had rarely opened the file in years suddenly had reasons to care what was in it.
Phase Two: Traffic Composition Flipped and Forced the Issue
The composition of crawler traffic tells the story clearly. A growing share of crawler requests now exists to pull down training data rather than to index pages for search. In practical terms, a large portion of what hits a site from a bot has nothing to do with being found in search. It’s feeding something else.
That changes the math on every allow and disallow line. A publisher who was happy to be crawled when it meant referrals has to decide whether the same openness still makes sense when it means training data for a competitor’s model. Many site owners landed on the same answer at once, and the file started changing in a lot of places.
Phase Three: Owners Split ‘Training’ From ‘Answering’
The subtler shift is that the same company often operates multiple crawlers with different jobs. OpenAI’s documentation spells out three separate user agents: one for training, one for surfacing results inside ChatGPT, and one for fetches a user kicks off by asking a question. Each is controlled independently in robots.txt. That gives site owners a lever they didn’t have a year ago.
The emerging pattern is to allow the answering bots and block the training ones. The logic is straightforward. Being cited in an AI answer at least carries the chance of attribution and a visit, while being used as training data gives nothing back.
Phase Four: The User-Agent Loophole Became the Real Fight
The real controversy is what happens when an AI agent fetches a page because a human asked it to. Some operators argue a user-prompted fetch is the user visiting, not a crawler, so the robots exclusion rules don’t apply. Others argue a bot is a bot no matter who sent it, and that bypassing the file is scraping with better manners.
This isn’t theoretical. A prominent infrastructure provider publicly called out one AI company in 2025 for fetching pages that had disallowed it, and the dispute surfaced a proposed cryptographic standard called Web Bot Auth for identifying agent requests at the HTTP layer.
Site owners who want a defensible position are no longer relying on robots.txt alone. They’re pairing it with firewall rules, bot verification, and in some cases paid-access gates for AI traffic specifically.
Phase Five: Write the File Like It’s a Policy, Not a Note to Yourself
The practical rewrite most site owners are doing right now has four moves. None of them are complicated. All of them matter more than they used to.
The rewrite isn’t about picking a side in the AI debate. It’s about catching up to the fact that a file written for one kind of visitor is now the gate for several kinds, each with different incentives. The owners doing this well treat robots.txt as a living policy. The ones who don’t will find the decisions made for them, one unannounced crawler at a time.See More
