Robots.txt 2026: managing AI crawler budgets for infrastructure leads

Robots.txt now decides which AI companies train on your pages and which cite them. Updated October 2026 with Google's AI Overviews opt-out, Cloudflare's 15 September changes and the agents that may ignore robots.txt.

Muhammad Zeeshan
Muhammad ZeeshanFounder & CEO
Updated October 7, 20265 min read
Minimalist 3D illustration of a futuristic server room managing an AI Crawler Budget. A central digital lock triages traffic: glowing red "Training Scrapers" are blocked by a red "X" to save bandwidth, while glowing green "Good Agents" are permitted with a checkmark to drive referral traffic. Labels at the bottom read "Blocked: Training Scrapers" and "Allowed: Good Agents
Share
Share

For most of its history, robots.txt told Googlebot where the sitemap was and which folders to skip. In 2026 it also decides which AI companies can train on your pages and which can quote them in an answer, and it states your terms to the agents that fetch pages for users. That makes it a cost decision as much as an SEO one.

If you are not managing your AI crawler budget, you are paying for bandwidth and origin CPU that trains someone else's model, while the agents that would cite you and send a visitor back get treated the same as the ones that never will.

Updated 2 October 2026. This version corrects what we said about Cloudflare's 15 September change, adds Google's new opt-out for AI Overviews and AI Mode, and adds the agents whose own makers say they may not follow robots.txt.

AI crawlers, October 2026

3
roles per vendor, not one
training, search, user-triggered
5
vendors whose user-triggered agents may skip robots.txt
by their own documentation
31 Aug
Google's AI Overviews opt-out went worldwide
a Search Console setting, not robots.txt
402
the status code for pay per crawl
Payment Required, not 403
Sources: crawler documentation from OpenAI, Anthropic, Google, Perplexity, Meta, Apple, Amazon and Mistral, plus Cloudflare's blog. Read 2 October 2026.

The engineering problem: the "shadow" crawl

Training crawlers such as GPTBot and CCBot collect pages in bulk to build datasets. They are not looking for today's change to answer today's question, so they tend to walk the whole site, miss CDN caches, hit endpoints that were never optimized and push P99 latency up. That is the shadow crawl: load that returns no visitors.

We are not going to put a percentage on it. The share varies widely from site to site, and the only number that matters is the one in your own logs. The last section shows how to get it in about ten minutes.

1. Triage: three roles, not two

The old advice split AI bots into "good agents" and "scrapers". The vendors now document three roles, usually with a separate user agent for each:

  • Training crawlers collect content to train models. Blocking them costs you no referral traffic.
  • Search and answer crawlers index pages so an AI search product can cite them. Block these and you leave that product's answers.
  • User-triggered fetchers load a page because a person asked an assistant to. Several vendors say these may not follow robots.txt at all, because a user, not a crawler, made the request.

Who runs which agent, October 2026

CompanyTrainingSearch and answersUser-triggered
OpenAIGPTBotOAI-SearchBotChatGPT-User (robots.txt may not apply)
AnthropicClaudeBotClaude-SearchBotClaude-User
GoogleGoogle-Extended (a control token, not a crawler)GooglebotGoogle-Agent (generally ignores robots.txt)
PerplexityNone documentedPerplexityBotPerplexity-User (generally ignores robots.txt)
MetaMeta-ExternalAgentMeta-WebIndexerMeta-ExternalFetcher (may bypass robots.txt)
AppleApplebot-Extended (a control token)ApplebotNone documented
AmazonAmazonbot (may be used for training)Amzn-SearchBotAmzn-User (may not follow all directives)
MistralMistralAI-TrainingMistralAI-IndexMistralAI-User
Common CrawlCCBotNoneNone

Sources, read 2 October 2026: OpenAI, Anthropic, Google's common crawlers and user-triggered fetchers, Perplexity, Meta, Apple, Amazon and Mistral.

Two corrections to the template this guide carried earlier in 2026:

  • FacebookBot is gone from Meta's documentation. Meta's training crawler is Meta-ExternalAgent, and Meta-WebIndexer feeds Meta AI search. A leftover FacebookBot rule does no harm. Do not block FacebookExternalHit: it draws the link preview when someone shares your page.
  • Anthropic documents three agents and no others. ClaudeBot, Claude-SearchBot and Claude-User are the only names on Anthropic's page, and Anthropic says its bots honor robots.txt. We had described Claude-Web as retired and anthropic-ai as a legacy name, but neither claim came from Anthropic, so both names are out of the template.

2. The 2026 robots.txt template

# Training crawlers: no referral value
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Meta-ExternalAgent
Disallow: /

User-agent: MistralAI-Training
Disallow: /

# Control tokens, not crawlers. Decide these rather than copy them:
# Google-Extended also takes you out of grounding in Gemini apps.
User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

# Search and answer crawlers: these produce citations
User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Meta-WebIndexer
Allow: /

User-agent: Amzn-SearchBot
Allow: /

User-agent: MistralAI-Index
Allow: /

# User-triggered fetchers: several may ignore this file anyway
User-agent: ChatGPT-User
Allow: /

User-agent: Claude-User
Allow: /

User-agent: Perplexity-User
Allow: /

User-agent: MistralAI-User
Allow: /

What the template leaves out, on purpose:

  • Googlebot. Blocking it removes you from Google Search, including AI Overviews and AI Mode. Google gives you a separate control for those, covered in the next section.
  • Amazonbot. Amazon says it may be used to train Amazon AI models, but it also serves Amazon products, and Amazon's Content Partners program rewards creators who allow it. Block it deliberately or not at all. Amazon also honors a noarchive robots meta tag as "do not use the page for model training".
  • OAI-AdsBot. If you advertise on ChatGPT, leave it alone. OpenAI uses it to check pages submitted as ads and says the data is not used to train models.

3. What Google lets you control now

Google-Extended is a control token, not a separate crawler. Google's documentation says it governs whether your content trains Gemini models and grounds answers in Gemini apps and Vertex AI, and that it "does not impact a site's inclusion in Google Search." Blocking it is a real trade: you leave Gemini's training data and Gemini's grounded answers, while Google Search is unaffected. Our earlier template described it only as an opt-out from Gemini training, which was half the story.

The Search Console control. Since 31 August 2026, every site can switch on "Exclude my site's links and content from Search generative AI features" in Search Console. It removes your pages from AI Overviews, AI Mode and Discover's AI features, and Google's help page is direct about the price: "You won't receive any traffic or impressions from these features." It does not affect training. For that, the same page points back to Google-Extended.

So there are now two separate Google decisions: whether your content trains models, and whether your pages appear in Search's AI answers. A business that wants to be cited should usually leave both open.

4. What Cloudflare changed on 15 September 2026

If your site sits behind Cloudflare, part of this decision may already have been made for you.

Correction. An earlier version of this guide said that from 15 September 2026, Cloudflare's default settings block "mixed-use" crawlers on any page carrying ads. That was the plan Cloudflare announced on 1 July 2026. What shipped is different: blocking mixed-use crawlers on ad pages is a setting the site owner chooses.

What did ship on 15 September 2026, per Cloudflare's announcement:

  • Separate settings for search, training and agents, including a "Disallow AI Training" option that writes a no-training preference into robots.txt.
  • Existing sites were moved over from their old "Block AI Bots" setting. If it was off, everything stays allowed. If it was on, training is now set to "Disallow AI Training" and agents are blocked on pages with ads.
  • New domains are offered presets, and sites that earn money from advertising are offered one that blocks agents on pages with ads.
  • Managed robots.txt is being retired in favor of Bot Preference Sync, which writes robots.txt from your dashboard settings.
  • Bing does not yet receive a no-training preference through robots.txt. Cloudflare says Microsoft is building support, targeted for early 2027.

Underneath it sits Cloudflare's Content Signals Policy, launched on 24 September 2025: plain-language lines in robots.txt that separate search, AI input and AI training. It went onto more than 3.8 million domains using Cloudflare's managed robots.txt, set to allow search and refuse training. These are stated preferences, not blocks.

Check what your CDN wrote before you write your own file. You may be editing on top of a default you never chose.

Can you charge crawlers instead of blocking them?

Yes, in a limited beta. Cloudflare's pay per crawl answers a request for a priced page with 402 Payment Required and a crawler-price header, instead of serving the page or returning a flat 403. Crawlers that want to pay have to sign their requests so Cloudflare can tell who they are. On 30 September 2026 Cloudflare added a Pay Per Use beta, which it presents as the next step after pay per crawl.

Whether either is worth doing depends on whether your content is scarce enough that a model builder would pay rather than skip it. For most business sites it is not, but the option exists.

5. Does robots.txt stop anyone?

Not on its own:

  • Undeclared crawlers. In August 2025 Cloudflare published evidence that Perplexity was using undeclared crawlers to get around no-crawl rules, and de-listed it as a verified bot in the same post. Perplexity disputed the findings.
  • User-triggered fetchers. OpenAI, Google, Perplexity, Amazon and Meta each say in their own documentation that agents acting on a user's request may not follow robots.txt.
  • The law. In December 2025 a US federal court ruled in Ziff Davis v. OpenAI that robots.txt is not a technical measure that controls access under the DMCA, comparing it to a sign asking visitors to keep off the grass. In the EU, the Commission's enforcement powers over general-purpose AI models took effect on 2 August 2026, and providers that signed its code of practice commit to following robots.txt.

Treat robots.txt as a statement of intent that well-behaved crawlers follow, and the edge as the place that enforces it. Signed requests are starting to close the identity gap: the IETF's Web Bot Auth working group published its first protocol draft in September 2026, and Google is testing the approach for Google-Agent.

6. Protecting the origin at the edge

Updating a text file is rarely enough. To manage the budget, put the same rules into your edge layer.

A web application firewall can reject a request before it reaches your application, so you do not pay the CPU cost of serving a crawler you have already decided to refuse. Keep the rules and robots.txt in step, ideally generated from one list, so the policy you state and the policy you enforce cannot drift apart.

This matters most for sensitive endpoints. If you have implemented API-First SEO to serve machine-readable data, only agents that can act on it should reach transactional endpoints, while training crawlers stay limited to public documentation. Pair this with nested JSON-LD for GraphRAG retrieval so retrieval agents get the facts they need from fewer pages.

7. Measuring the ROI of a crawler

Treat each AI crawler as a cost against a benefit:

  • Cost: total requests per month multiplied by your egress price, plus origin CPU load during crawl spikes.
  • Benefit: referral traffic from AI search, plus conversions you can attribute to agents.

A crawler with high cost and no benefit is a candidate for a block. An agent that hits your Action Schema endpoints to complete a purchase is spending your crawl budget on something that pays.

Audit your own crawl before you copy anyone's template

Every number in a post like this one, including ours, describes someone else's site. The only split that matters is yours, and it takes about ten minutes to get.

  1. Pull 30 days of edge or access logs and group requests by user agent. Most CDNs will do this in the dashboard.
  2. Bucket them into the three roles: training, search and answers, and user-triggered. The table above tells you which is which.
  3. Join to referrals. Look at how much traffic each bucket sent back over the same window. A crawler with cost and no referral is a candidate for a block; one feeding an answer engine that cites you is not.
  4. Check what your CDN already decided. If your robots.txt is managed or synced, read it before you write a new one.

Do that before adopting any template, including the one above. A block that removes you from an answer engine your buyers use costs more than a month of wasted bandwidth.

Conclusion

Letting every AI crawler take what it wants is no longer the default, and it should not be yours. Decide per role: refuse training if you choose, stay open to the crawlers that cite you, and enforce what matters at the edge, because the agents acting for users may not read robots.txt at all.

The next problem after rationing access is verifying it: a user-agent string is a claim, not proof. Where an operator publishes IP ranges or a reverse DNS pattern for its crawlers, check requests against them at the edge, and treat traffic that claims a known name but fails the check as unknown.

What changed in this guide

  • January 2026: first published.
  • Q2 2026: added Anthropic's three agents and Meta's training crawler.
  • 27 September 2026: removed illustrative bandwidth percentages from the text, and added Cloudflare's Content Signals Policy and pay per crawl.
  • 2 October 2026:
    • Corrected the Cloudflare 15 September section.
    • Added Google's Search Console opt-out and the full scope of Google-Extended.
    • Added the user-triggered fetchers that may ignore robots.txt, and Meta's, Amazon's and Mistral's current agents.
    • Rebuilt the template, and removed the diagram that still carried the old percentages.
  • 7 October 2026: rewrote the closing note on verifying crawlers.

Let's discuss it over a call.

Key takeaways

  • Block training crawlers (GPTBot, ClaudeBot, CCBot, Meta-ExternalAgent, MistralAI-Training). Allow search and answer crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot, Meta-WebIndexer).
  • Treat Google-Extended and Applebot-Extended as decisions, not defaults: Google-Extended also takes you out of grounding in Gemini apps.
  • Google's Search Console opt-out removes you from AI Overviews and AI Mode, and with it all traffic from those features.
  • On Cloudflare, read what Bot Preference Sync wrote into robots.txt before writing your own. Blocking agents on ad pages is a setting you choose.
  • Audit 30 days of logs by crawler role and referrals before copying any template, including ours.
Muhammad Zeeshan

Written by

Muhammad Zeeshan

Founder & CEO

Muhammad Zeeshan is a website, design, ecommerce, SEO, and digital growth specialist with 9+ years of experience helping clients build and improve their online presence. His work covers Webflow, WordPress, Shopify, Framer, Figma, SEO, AEO, GEO, social media, presentations, and website support.

Keep reading

Related articles.

More on the same thread, picked by tag and category, not chronology.

9 min read

The AEO Audit Checklist

An interactive AEO audit with a weak-versus-strong example for every item and a live self-scoring widget. Grade your site in five minutes.

Muhammad ZeeshanMuhammad Zeeshan
Read

Newsletter

New guides, straight to your inbox.

Practical notes on websites, SEO, AEO, GEO, ecommerce, and automation, sent when new guides are published.

No spam. Unsubscribe any time.

Ready when you are

Want help with AEO & GEO?

Send us your website and your goal. We will tell you honestly what should be done first, whether or not we work together.