How to Block AI Crawlers: The Complete GPTBot and robots.txt Reference

Blocking AI crawlers takes two lines in your robots.txt. Deciding whether you should is the hard part. Here’s the short version: GPTBot collects training data, so blocking it protects your content but wins you nothing in ChatGPT. OAI-SearchBot is what actually surfaces you in ChatGPT answers. Block the wrong one and you delete yourself from the fastest-growing referral channel on the web.

Most guides hand you a copy-paste blocklist and stop. That’s not enough. Robots.txt is a request, not a fence. Some crawlers ignore it. Some don’t identify themselves at all. And the crawler hitting your server may not be the thing deciding whether you get cited.

This page is the technical reference: every major AI user agent, the exact syntax, how to verify a crawler is real, how to enforce at the CDN, and the 2026 data on what blocking actually does to your visibility.

What GPTBot actually is (and what it isn’t)

OpenAI runs four separate crawlers. They have different jobs, different IP ranges, and different consequences when you block them. Treating them as one thing is the single most common mistake in AI crawler blocking. Each has its own published IP range file and its own robots.txt token, so you can allow one and refuse another without compromise.

According to OpenAI’s official bot documentation, each one is controlled independently in robots.txt, and changes take roughly 24 hours to propagate through OpenAI’s systems.

User agent tokenWhat it doesBlocking it meansIP list
GPTBotCrawls pages that may be used to train OpenAI’s foundation modelsYour content isn’t used for model training. No effect on ChatGPT search citations.openai.com/gptbot.json
OAI-SearchBotBuilds the index behind ChatGPT’s search featureYou disappear from ChatGPT search results and citations.openai.com/searchbot.json
ChatGPT-UserFetches a page live when a user or a Custom GPT asks for itChatGPT can’t open your page on request. OpenAI notes robots.txt “may not apply” here, since a human triggered it.openai.com/chatgpt-user.json
OAI-AdsBotChecks landing pages submitted as ChatGPT ads for safety and relevanceYour landing pages can’t be validated for ChatGPT advertising. Not used for training.openai.com/adsbot.json

The user agent strings you’ll see in logs

These are the real strings OpenAI publishes. Note that the training crawler doesn’t pretend to be a browser, while OAI-SearchBot carries a full Chrome fingerprint:

  • Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot
  • Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot
  • Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot
  • Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-AdsBot/1.0; +https://openai.com/adsbot

Version numbers change. Match on the token, never the whole string. OpenAI also ships a variant that appends a robots.txt marker to the user agent when the request is fetching your rules file rather than a content page, so don’t count those as page crawls.

Why the distinction decides everything

Training and retrieval are separate pipelines. A model trained in 2024 doesn’t know your 2026 pricing page. When ChatGPT answers a question about your product today, it usually runs a search and reads live pages. That’s OAI-SearchBot and ChatGPT-User doing the work, not the training crawler.

So the trade-off is cleaner than people assume. Blocking GPTBot costs you almost nothing in the short term. Blocking the retrieval crawlers costs you the citation.

The AI crawler reference table

Here’s every crawler worth a rule in 2026, grouped by operator. “Respects robots.txt” reflects documented policy and observed behaviour, which aren’t always the same thing.

User agentOperatorPurposeRespects robots.txt
GPTBotOpenAIFoundation model trainingYes
OAI-SearchBotOpenAIChatGPT search indexYes
ChatGPT-UserOpenAILive user-triggered fetchPartly — OpenAI says rules “may not apply”
OAI-AdsBotOpenAIAd landing page validationYes
ClaudeBotAnthropicModel trainingYes
Claude-SearchBotAnthropicClaude search qualityYes
Claude-UserAnthropicLive fetch for Claude answersYes
Google-ExtendedGoogleOpt-out token for Gemini training. Does not crawl.Yes (it’s a directive, not a fetcher)
GooglebotGoogleSearch index, AI Overviews, AI ModeYes — but blocking it removes you from Search
Google-CloudVertexBotGoogleVertex AI agent grounding for customersYes
PerplexityBotPerplexityIndex for Perplexity answersYes (declared)
Perplexity-UserPerplexityLive fetch on user requestPartly — treated as user-initiated
ApplebotAppleSiri and Spotlight search indexYes
Applebot-ExtendedAppleOpt-out token for Apple Intelligence training. Does not crawl.Yes (directive only)
Meta-ExternalAgentMetaTraining data for Meta AIYes
Meta-ExternalFetcherMetaLink fetch for assistant answersPartly
AmazonbotAmazonAlexa and Amazon AI productsYes
CCBotCommon CrawlOpen dataset used by many model buildersYes
BytespiderByteDanceTraining data for ByteDance modelsPoor track record
DuckAssistBotDuckDuckGoAI-assisted answersYes
MistralAI-UserMistralLe Chat live browsingYes
cohere-aiCohereModel training and retrievalYes

Two of these aren’t crawlers at all

Google-Extended and Applebot-Extended never request a page. They’re permission tokens. Apple states plainly that “Applebot-Extended does not crawl webpages,” and that pages disallowing it still appear in search results. Google-Extended works the same way: it governs whether content already fetched by Googlebot may be used to train Gemini.

That has a sharp consequence. You cannot use Google-Extended to keep yourself out of AI Overviews. AI Overviews are built on the Search index, which means Googlebot. Since June 2, 2026, Google has been rolling out a separate opt-out toggle in Search Console that removes a site from AI Overviews and AI Mode without affecting ordinary Search rankings. It started with a subset of UK site owners.

Robots.txt syntax that actually works

Robots.txt lives at the root of every host and protocol you serve: https://example.com/robots.txt. Subdomains need their own file. A few rules trip people up constantly:

  • User agent tokens are matched case-insensitively, but write them exactly as documented anyway.
  • A crawler obeys only the most specific matching group. Name a bot in its own block and your User-agent: * rules are ignored for it entirely.
  • Disallow: / blocks everything. Disallow: with nothing after it allows everything.
  • Blank lines separate groups. A stray blank line inside a group silently splits it.
  • Robots.txt stops crawling, not indexing or linking. It is not a security control.

Recipe 1: block training, keep every citation

This is the right default for most brands. You stay retrievable in ChatGPT, Claude, Perplexity and Google, but your content doesn’t feed model training. The explicit Allow: / blocks are technically redundant, since anything not disallowed is allowed, but they make your intent obvious to whoever edits this file next.

# Training crawlers - blocked
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: Meta-ExternalAgent
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Bytespider
Disallow: /

User-agent: Amazonbot
Disallow: /

# Retrieval and search crawlers - allowed
User-agent: OAI-SearchBot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: Claude-User
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Perplexity-User
Allow: /

User-agent: Applebot
Allow: /

User-agent: Googlebot
Allow: /

Recipe 2: hard block on everything AI

For publishers who want leverage in a licensing negotiation, or who simply want out. Note this removes you from ChatGPT and Perplexity citations too.

User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: OAI-AdsBot
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: Meta-ExternalAgent
User-agent: Meta-ExternalFetcher
User-agent: Amazonbot
User-agent: CCBot
User-agent: Bytespider
User-agent: DuckAssistBot
User-agent: MistralAI-User
User-agent: cohere-ai
Disallow: /

Stacking multiple User-agent lines above one Disallow is valid and applies the rule to all of them. Keep Googlebot out of this list unless you genuinely want to leave Google Search.

Recipe 3: protect the crown jewels only

Path-level blocking beats site-level blocking when your moat is a specific dataset, tool or gated library. Let the training crawler have your blog. Keep it away from the thing competitors would love to have modelled.

User-agent: GPTBot
Disallow: /research/
Disallow: /data/
Disallow: /members/
Allow: /

User-agent: ClaudeBot
Disallow: /research/
Disallow: /data/
Disallow: /members/
Allow: /

One caution on Disallow: /members/: you’ve just published the path you consider valuable. If it’s genuinely sensitive, put it behind authentication instead.

Pair whatever you choose with a clear llms.txt file so the crawlers you do welcome can find your best pages fast. Our llms.txt generator will build one from your sitemap.

How to verify a real GPTBot request by IP

Anyone can send a request claiming to be an OpenAI crawler. Scrapers do it constantly, because it slips past naive rate limits and it poisons your analytics. Verification is the only reliable answer, and it takes three steps.

  1. Pull the official IP list. OpenAI publishes a JSON file per bot. https://openai.com/gptbot.json currently lists 21 IPv4 CIDR blocks, including ranges like 132.196.86.0/24 and 20.125.66.80/28. The file carries a creationTime field so you can tell how fresh your copy is. Anthropic publishes the equivalent at https://claude.com/crawling/bots.json.
  2. Match the request IP against the CIDR ranges. In Python, ipaddress.ip_address(ip) in ipaddress.ip_network(cidr) is all you need. Cache the JSON and refresh it daily. Never hardcode individual addresses.
  3. Fall back to reverse DNS for Google. Google doesn’t publish a static list for every crawler, so use the classic method: reverse-lookup the IP, confirm the hostname ends in googlebot.com or google.com, then forward-lookup that hostname and check it resolves back to the same IP. A single reverse lookup is not enough, because reverse DNS records can be forged.

The newer method: signed requests

Reverse DNS and IP lists don’t scale to thousands of AI agents. The emerging alternative is Web Bot Auth, an IETF draft that has crawlers sign requests with an HTTP Message Signature tied to a published public key. Cloudflare has already shipped support for cryptographically verified bots. If you’re building verification today, leave room for a signature header path alongside your IP checks.

Why this matters: a spoofed request that passes your check gets recorded as legitimate AI crawler traffic, and you’ll make strategy decisions on a number that isn’t real. A meaningful share of traffic claiming to be a major AI crawler comes from IPs the operator has never owned.

Blocking at the CDN and WAF level

Robots.txt is voluntary. If a crawler ignores it, you need enforcement at the edge. Tollbit’s Q2 2025 data found 13.26% of AI bot requests ignored robots.txt directives, up from 3.3% in Q4 2024. That trend is why edge rules stopped being optional.

The escalation ladder, cheapest to strongest:

  1. Robots.txt. Free, instant, honoured by every major operator. Start here.
  2. Meta tags and headers. <meta name='robots' content='noai, noimageai'> is widely used but not universally honoured. X-Robots-Tag in the HTTP response works for non-HTML files like PDFs.
  3. User agent blocking at the server. An nginx rule such as if ($http_user_agent ~* '(GPTBot|ClaudeBot|CCBot|Bytespider)') { return 403; } stops the polite ones cold. Trivial to evade by spoofing.
  4. IP and ASN blocking. Block the published ranges, or the whole ASN for repeat offenders. Effective, but heavy-handed, and it breaks when ranges change.
  5. Managed bot rules at the CDN. Cloudflare, Fastly and Akamai all ship AI bot categories now. This is behavioural detection, so it catches undeclared crawlers that user agent rules miss.

Return 403, not 404

Serve a real 403 Forbidden to blocked crawlers. A 404 tells the operator your page doesn’t exist, which can strip you from indexes you wanted to stay in. A 429 with a Retry-After header is the gentler option when your problem is crawl volume rather than consent.

And rate-limit before you block. Many “AI crawlers are destroying my server” complaints turn out to be one aggressive bot hammering faceted URLs. A 5-requests-per-second cap on /search? and filter parameters solves it without giving up any visibility.

Don’t break your own edge rules

Two failure modes come up constantly. The first is a regex that matches too broadly: a rule targeting bot will also catch Googlebot, Bingbot and every uptime monitor you run. Anchor your patterns to full tokens. The second is blocking at the CDN while leaving robots.txt permissive, which produces a confusing state where crawlers keep requesting pages they’ll never receive and your logs fill with 403s that look like an outage.

Test every change from outside your network. A single curl with the crawler’s user agent and the -I flag tells you immediately whether the rule fires and what status code it returns.

Cloudflare’s AI crawler controls and pay-per-crawl

Cloudflare sits in front of a large share of the web, so its defaults are effectively industry policy. Two announcements reshaped this space.

On July 1, 2025, Cloudflare began blocking AI crawlers by default for new domains and launched Pay Per Crawl in private beta. On July 1, 2026, it went further, splitting AI traffic into three categories site owners control separately: Search, Agent and Training. From September 15, 2026, Training and Agent bots are blocked by default on ad-monetised pages for new domains, while Search stays allowed.

That three-way split matters because it forces operators to declare intent. A crawler that refuses to separate its training fetches from its search fetches gets treated as training, and blocked.

How pay-per-crawl actually works

The mechanism revives an HTTP status code almost nobody used. When a crawler requests a priced page, Cloudflare returns HTTP 402 Payment Required with a crawler-price header. The crawler can retry with crawler-exact-price to accept, or send crawler-max-price up front to declare a ceiling. On success it gets a 200 and a crawler-charged confirmation. Cloudflare acts as merchant of record.

It’s elegant. It’s also only meaningful if the AI companies choose to participate, and adoption has been slow. Treat it as a signalling mechanism and a data source rather than a revenue line, unless you’re a large publisher.

Cloudflare’s AI Crawl Control dashboard is worth turning on regardless. It tracks which crawlers are violating your robots.txt directives, which is the fastest way to learn whether your rules are honoured at all. The July 2026 update also added an optional fourth field to robots.txt for declaring how content may be used rather than just who may fetch it: use=immediate, use=reference (the default) and use=full. It carries no enforcement yet, but it’s the first serious attempt to express intent in a file designed for access control.

The Perplexity precedent

On August 4, 2025, Cloudflare published evidence that Perplexity was using undeclared crawlers with a generic Chrome user agent, rotating IPs outside its published ranges and across different ASNs, generating an estimated 3 to 6 million requests per day on top of its declared bot. Cloudflare removed Perplexity from its Verified Bots programme and shipped detection signatures to all customers.

The lesson isn’t about one company. It’s that a user agent string is a claim, not a fact, and any blocking strategy built purely on user agent matching has a hole in it.

Does blocking cost you citations? The data

This is where most advice is guesswork. There’s now real data, and it’s counterintuitive.

A BuzzStream study published March 19, 2026 analysed 4 million AI citations from 3,600 prompts across ChatGPT, Gemini, Google AI Overviews and AI Mode, focusing on the top 50 news sites and what they blocked. The findings:

  • 88.2% of sites blocking OpenAI’s training crawler were cited anyway.
  • 92.3% of sites blocking Google-Extended still appeared in citations.
  • 82.4% of sites blocking OAI-SearchBot still showed up.
  • 70.6% of sites blocking ChatGPT-User were still cited.
  • CNBC appeared 1,298 times while blocking three crawlers. Yahoo appeared in nearly 30,000 citations while blocking Google-Extended.

Why? Because models cite what they know from third-party coverage, aggregators, syndication and Wikipedia. Blocking removes your page from the fetch, not your brand from the model’s world.

The honest read: blocking is far less catastrophic for citations than vendors claim, and far less protective of your content than publishers hope. You lose the click, not the mention.

The traffic side of the ledger

Cloudflare’s crawl-to-refer data quantifies the asymmetry. In the first week of August 2025, Anthropic’s ratio was roughly 50,000 crawls per referral, OpenAI’s about 887:1 and Perplexity’s about 118:1. Roughly 80% of all AI bot crawling in July 2025 was training traffic.

Those numbers explain publisher anger. They also explain why “just block everything” is a defensible position for a subscription news business and a bad one for a SaaS company that needs discovery.

How many sites actually block?

Fewer than the discourse suggests, though it’s climbing fast. Ahrefs analysed roughly 140 million websites in a study published May 21, 2025 and found 5.89% blocked GPTBot. A June 2026 analysis of the top 10,000 domains found 17.6% blocking GPTBot but only 6.6% blocking OAI-SearchBot — sites are close to three times more willing to be cited than to be trained on. Among news publishers specifically, roughly half block the training crawler.

The Register reported on December 8, 2025 that the number of sites blocking GPTBot had risen from 3.3 million in July 2025 to 5.6 million, a 70% jump in five months.

How to read your server logs for AI crawlers

Before you change anything, find out what’s actually hitting you. Most people are surprised.

The fastest grep against a standard combined-format access log:

grep -Ei 'GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-User|PerplexityBot|Bytespider|CCBot|Meta-ExternalAgent|Amazonbot|Applebot' access.log \
  | awk '{print $1}' | sort | uniq -c | sort -rn | head -20

Four things to look for:

  1. Volume vs. value. Divide crawler hits by referral sessions from that platform. If a training crawler pulls 40,000 pages a month and ChatGPT sends you nine visits, you now know your own crawl-to-refer ratio.
  2. Status codes. A wall of 404s or 403s means the crawler is finding broken paths or being blocked by a rule you forgot about. Both waste your crawl budget and theirs.
  3. Which URLs. If crawlers burn requests on faceted search and tag archives instead of your money pages, that’s an information architecture problem, not a blocking problem.
  4. Unverified claims. Any request claiming to be a known crawler from an IP outside the published ranges is a spoof. Count those separately.

Turn the log into a decision

Logs tell you who’s crawling. They don’t tell you who’s citing you. For that, run a prompt-level check — our guide on how to check if ChatGPT cites your site walks through the method, and the AI visibility checker automates it across engines.

One caveat on analytics: most AI crawler traffic never appears in Google Analytics, because these bots don’t execute the JavaScript that fires your tag. Server logs, or your CDN’s own analytics, are the only complete picture. If you’ve been judging AI crawler impact from GA4, you’ve been looking at roughly nothing.

The decision framework: publisher or brand?

There’s no universal right answer, but there are two clear archetypes and they point in opposite directions.

If your content is the product

News, research, paid databases, courses, anything behind a paywall or funded by ads. Your content has direct licensing value and AI answers cannibalise your traffic. Block the training crawlers, enforce at the edge, and use the block as leverage. Publishers who blocked early ended up with licensing deals. Publishers who left the door open ended up with neither payment nor traffic.

Practical stance: block GPTBot, ClaudeBot, CCBot, Bytespider, Meta-ExternalAgent and Google-Extended. Decide on search crawlers based on whether AI referrals convert for you. Turn on your CDN’s controls. Track everything.

If content is how you get found

SaaS, agencies, ecommerce, B2B, local services — most of the market. Your blog is a marketing cost, not an asset you license. Being absent from AI answers is a real loss, and being trained on is a theoretical one. Allow the retrieval crawlers without hesitation.

The training question here is close to a coin flip. Training exposure has some brand value: models that have read your documentation describe your product more accurately. But it produces no measurable traffic. Blocking GPTBot is a defensible, low-cost hedge. Blocking OAI-SearchBot is not.

The five questions that settle it

  1. Does anyone pay for this content directly? Yes means block training. No means the licensing argument doesn’t apply to you.
  2. Are AI referrals converting? Check analytics for chatgpt.com, perplexity.ai and claude.ai referrers. If they convert above your organic average, protect that channel at all costs.
  3. Is crawler load costing you real money? Rate-limit first. Only block if limits don’t fix it.
  4. Would a competitor benefit from a model trained on this? If yes, block the training crawlers on those paths specifically.
  5. Can you tell the difference in six months? Baseline your crawl and citation data now, or you’ll never know whether the change worked.

Whichever way you go, document the decision and revisit it quarterly. This landscape changed three times in the last eighteen months.

What to do after you’ve decided

Blocking is a defensive move. If visibility is the goal, the offensive work matters more: clean structure, extractable answers, entity clarity and citable claims. Start with a GEO audit to see how retrievable your pages actually are, then work through the fundamentals in our LLM SEO guide.

Robots.txt decides whether the door is open. It doesn’t decide whether anything inside is worth quoting.

Frequently asked questions

Should I block GPTBot?

For most brands there’s no strong reason either way, and blocking OAI-SearchBot instead is a clear mistake. The training crawler produces no measurable traffic, though it does shape how accurately models describe your product. Publishers who license content should block it. Marketers chasing visibility should leave the retrieval crawlers alone.

Does blocking GPTBot stop ChatGPT from citing my site?

No. It handles training only, while OAI-SearchBot and ChatGPT-User handle live retrieval and citations. A March 2026 BuzzStream study of 4 million citations found 88.2% of blocking sites were still cited anyway. To leave ChatGPT answers entirely you’d have to block the search and user agents too.

Does blocking AI crawlers hurt my Google rankings?

No, as long as you don’t block Googlebot. Google-Extended and Applebot-Extended are training opt-out tokens that never crawl, and Google confirms they don’t affect Search. The one real risk is a badly scoped CDN rule that catches Googlebot alongside the AI bots.

How do I stop appearing in Google AI Overviews?

Google-Extended won’t do it, because AI Overviews run on the Search index. Since June 2, 2026 Google has been rolling out a Search Console toggle that removes a site from AI Overviews and AI Mode without affecting regular Search rankings or Discover. It started with a subset of UK site owners. The older workaround is nosnippet or a restrictive max-snippet, which also strips your normal search snippets.

Do AI crawlers actually respect robots.txt?

The major ones mostly do. Tollbit measured 13.26% of AI bot requests ignoring robots.txt in Q2 2025, up from 3.3% in Q4 2024. Bytespider has a poor record, and Cloudflare documented Perplexity using undeclared crawlers in August 2025. Edge enforcement is the answer for anything that matters.

How do I verify a request is really from an OpenAI crawler?

Match the source IP against OpenAI’s published CIDR ranges, which for the training bot live at openai.com/gptbot.json and currently list 21 IPv4 blocks. Never trust the user agent string alone, since it’s trivially spoofed. Refresh the JSON daily rather than hardcoding IPs, and use forward-confirmed reverse DNS for Google’s crawlers.

What is the difference between GPTBot and ChatGPT-User?

The first crawls broadly to gather training data on OpenAI’s own schedule. ChatGPT-User fetches one page at a time because a human asked ChatGPT to look at it. OpenAI states that robots.txt rules may not apply to user-initiated requests, so block that agent at the edge if you need it genuinely stopped.

How long does a robots.txt change take to work?

OpenAI says roughly 24 hours for its systems to pick up changes. Other operators cache robots.txt for anywhere from a few hours to a week. If you need something blocked immediately, use a server or WAF rule and let robots.txt catch up.

Will blocking AI crawlers reduce my server load?

Usually yes, and sometimes dramatically. Around 80% of AI bot crawling in July 2025 was training traffic, the most volume-heavy and least valuable category. Try rate-limiting first, since aggressive crawling often targets faceted URLs and tag archives rather than real content. Check the traffic is verified rather than spoofed before you act.

Is there a robots.txt directive for AI training specifically?

Not a universal one yet. Google-Extended and Applebot-Extended are vendor-specific opt-out tokens. Cloudflare added an optional content-signal field to robots.txt in July 2026 with use=immediate, use=reference and use=full values, but adoption is early. For now, name each crawler explicitly.

What is pay-per-crawl and should I enable it?

Cloudflare’s pay-per-crawl returns HTTP 402 with a crawler-price header when a bot requests a priced page, then charges via crawler-exact-price on retry. It launched in private beta on July 1, 2025. It’s worth enabling for the data if you’re already on Cloudflare, but treat revenue as speculative until more AI companies opt in.

Should I use llms.txt instead of robots.txt?

They do different jobs and neither replaces the other. Robots.txt controls access; llms.txt is a curated map pointing AI systems at your most useful pages. Use robots.txt for permission decisions and llms.txt to help the crawlers you’ve allowed find your best content quickly.

zulqarnain, founder of LLM Optimization

Written by

zulqarnain

Writes about how AI search engines such as ChatGPT, Google AI Overviews, Perplexity, Gemini and Claude choose the sources they cite.

Scroll to Top