Short answer: AI crawlers come in two kinds. Search crawlers fetch pages to answer a person’s question, so blocking them stops AI from reading, citing or recommending you. Training crawlers collect pages to teach future models, so blocking them is a legitimate choice with a longer-term cost. Set the rules per crawler in robots.txt, then check that your firewall or CDN isn’t blocking the search crawlers anyway. Unvibe reads robots.txt rule by rule and visits your site as each crawler.
Two kinds of AI crawler
Here is how Unvibe sorts the crawlers it knows about. Check each company’s own documentation for the current purpose of its crawlers, because names and roles change.
| Crawler | Operator | Kind | What it is used for |
|---|---|---|---|
| OAI-SearchBot | OpenAI | Search | ChatGPT search results |
| ChatGPT-User | OpenAI | Search | ChatGPT when a person asks it to open a page |
| PerplexityBot | Perplexity | Search | Perplexity answers |
| Claude-SearchBot | Anthropic | Search | Claude search results |
| GPTBot | OpenAI | Training | Training OpenAI models |
| ClaudeBot | Anthropic | Training | Training Anthropic models |
| Google-Extended | Training | Gemini training and grounding. A robots.txt token, separate from Googlebot, which crawls for Google Search |
Two doors stand between those crawlers and your pages, and they are separate. robots.txt is a file that asks crawlers to stay out of certain paths. A firewall or CDN is the thing that actually serves, challenges or refuses a visit. A site can say “welcome” in robots.txt and still turn the same crawler away at the firewall.
Pick a policy before you touch a file
- Open to all. Nothing blocked. AI can read you, cite you, and future models can learn from you first-hand.
- Search in, training out. AI search can read and recommend you. Future models are not trained on your pages by the crawlers you block. The cost is that what AI says from memory will rest on what other sites say about you, since it learns nothing first-hand.
- Closed. Both blocked. AI search stops mentioning you, so only choose this if you don’t want AI traffic at all.
A business that wants customers to find it through AI assistants will want one of the first two. The third is often an accident, which is the next section.
Four ways it goes wrong
1. A blanket rule you didn’t write
A robots.txt copied from a staging setup or a template can block everything. It looks like this, and it blocks every crawler that obeys it, search crawlers included:
User-agent: * Disallow: /
2. A noindex left behind
Builders add <meta name="robots" content="noindex"> while a site is “private”, and it is easy to forget when you go live. It tells search engines and AI not to index the page, and the same instruction can also arrive as an X-Robots-Tag response header.
3. One switch that blocks all “AI bots”
Some CDNs, Cloudflare among them, have a setting that blocks AI bots as a group, search and training alike. If you only meant to stop training, it also stops the crawlers that let AI answer questions about you.
4. robots.txt says yes, the firewall says no
Bot protection can challenge or block a crawler that robots.txt welcomes. Then ChatGPT, Claude and Perplexity can’t open your site at all. Our own site had this: a Cloudflare setting was serving AI crawlers a robots.txt that turned them away, and switching it off took unvibe.app from 93 to 100. The story is in the case study.
The robots.txt for “search in, training out”
A crawler follows the rules in the group that names it, and falls back to the * group if none does. So you name the training crawlers you want to block, and leave everything else open. The domain below is a placeholder, so use your own.
# AI training crawlers: blocked User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Google-Extended Disallow: / # Everyone else, including AI search crawlers and Googlebot: welcome User-agent: * Allow: / Sitemap: https://www.example.com/sitemap.xml
For “open to all”, keep only the last group and the sitemap line. Keep the Sitemap line either way: it is how crawlers find your list of pages.
robots.txt is a request, not a lock. Well-behaved crawlers follow it. Anything that needs to be kept private shouldn’t be on a public page in the first place.
The firewall side
After robots.txt, look at the place that serves your site: the host, a CDN such as Cloudflare, a firewall or a security plugin.
- Find the bot protection or “AI crawlers” settings, including any option that blocks AI bots and any managed robots.txt the CDN adds for you.
- If you want AI search, make sure OAI-SearchBot, ChatGPT-User, PerplexityBot and Claude-SearchBot are allowed. Most tools have a “verified bots” or “AI crawlers” switch.
- If you block training on purpose, do it in robots.txt, where each crawler is named, rather than with a switch that blocks everything.
Then test it by asking for your homepage with a crawler’s user agent:
curl -sI -A "Mozilla/5.0 (compatible; GPTBot/1.2; +https://openai.com/gptbot)" https://www.example.com/
You want a 200 and the page. One caveat: some firewalls only let a crawler through when the visit comes from that crawler’s own network, so a request from your laptop can be refused even though the real crawler gets in. A refusal here is a reason to look, not proof of a block.
Where Unvibe fits
You can do all of this by hand. These are the parts Unvibe automates.
Can the crawlers get in?
- Reads robots.txt rule by rule for each AI crawler, and shows the line that blocks it.
- Visits your homepage as GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and Googlebot with their real user agents, to catch a CDN, firewall or bot protection that turns them away while browsers get the page.
- Treats the two kinds differently. A block on a search crawler is flagged as a problem, because AI search then stops mentioning you. A block on training crawlers is flagged as an opportunity, with the advice to keep it if it is deliberate.
- Says when it can’t tell. If a site turns away every visit that claims to be a crawler, even Googlebot, the report says it couldn’t be tested instead of guessing.
- Verifies the fix: after you change robots.txt or CDN settings, run a checkup. It re-reads the live robots.txt and visits as the crawlers again.
One honest trade-off: the robots.txt check in the score passes when AI crawlers such as GPTBot and ClaudeBot are let in, so a deliberate training block can show as a failing check. Decide your policy first; the score follows it, not the other way round.
The checks are described in the checkup docs, and the limits in the methodology. If AI can read your site but still doesn’t know what you are, the next step is saying it plainly.
Questions
If I block GPTBot, will ChatGPT stop mentioning my site?
GPTBot is the crawler that collects pages to train OpenAI models, while OAI-SearchBot and ChatGPT-User are the ones behind ChatGPT search and opening pages. Blocking only GPTBot leaves live search able to read you, but check the operator’s own documentation, because crawler roles can change.
Is robots.txt a lock?
No. It is a request that well-behaved crawlers follow. A firewall or CDN is what actually turns visitors away, which is why the two can disagree.
Will blocking training crawlers lower my Unvibe score?
It can. The robots.txt check looks at crawlers such as GPTBot and ClaudeBot, so a deliberate block may show as a failing check. The audit labels it an opportunity and says to keep the block if it is deliberate.