All posts
· The Topic Modeler Team

GPTBot, ClaudeBot, PerplexityBot: Should You Let AI Crawlers In?

GPTBot, ClaudeBot, PerplexityBot: blocking them erases you from AI answers. Get the decision framework and examples from Topic Modeler.

GPTBot, ClaudeBot, PerplexityBot: Should You Let AI Crawlers In?

We have this conversation with clients constantly, and it arrives from two opposite directions. Some clients are hyperprotective: worried about customer data, or convinced competitors are mining their content for ideas. Others have the opposite problem and do not know it; somewhere along the way a developer blocked every bot "to be safe," and the business quietly vanished from AI answers without anyone connecting the two events.

Both groups are asking the same question: should we let AI crawlers in? Our answer, in almost every case, is yes for the crawlers that put you in front of customers, and a genuine judgment call for the rest. Here is the framework we walk clients through.

Know who is knocking

The major AI crawlers, and what each one feeds:

Crawler Operator What it feeds
GPTBot OpenAI Model training
OAI-SearchBot OpenAI ChatGPT search (live retrieval)
ClaudeBot Anthropic Model training
PerplexityBot Perplexity Live answer retrieval
Google-Extended Google Gemini training (blocking it does not affect Google Search)

The column that matters is the last one, and the distinction that matters is training versus live retrieval. Training bots collect content that shapes what future models know. Retrieval bots fetch content right now to answer a live question, with a citation and often a link. Blocking a retrieval bot removes you from the answers your prospects are reading today. That is the tradeoff in one sentence: blocking protects content, and it also erases you.

Our verdict for businesses that want to be found

Let retrieval bots in, decide on training bots case by case. If AI answers are becoming a discovery channel for your category (they are), retrieval access is not optional; it is the trust and citation pipeline working as designed. Training access is a philosophical and business call: some clients like the idea of future models knowing their brand deeply, others object on principle. Either choice is defensible. Blocking retrieval by accident is not.

What we do urge everyone to keep crawlable is the whole authority surface: your content, your structure, and the machine-readable files like llms.txt and JSON-LD that describe it. An invisible authority is not an authority.

What to block regardless

Letting bots in does not mean letting them everywhere. We typically limit Disallow rules to pages that would never help you in any index:

  • Admin and login pages (/wp-admin/, /dashboard/) with no search value
  • Internal search result pages, which generate infinite low-quality URLs
  • Thank-you and confirmation pages you never want appearing anywhere

Copy-paste starting points

Welcome retrieval, block training, protect the utility pages:

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: *
Disallow: /wp-admin/
Disallow: /search/
Disallow: /thank-you/

Or welcome everything except the utility pages:

User-agent: *
Disallow: /wp-admin/
Disallow: /search/
Disallow: /thank-you/

Note what is absent from both: no Disallow for OAI-SearchBot or PerplexityBot. Those are the bots carrying your name into answers.

Make sure the open door leads somewhere

One last thing we tell every client after the robots.txt is settled: access is the precondition, not the payoff. A crawler you let in still has to find deep coverage, clean structure, and connected expertise, or it leaves with nothing worth citing.

That is the part Topic Modeler shows you. Map what the crawlers will actually encounter on your site, find the gaps before they do, and make the visit worth it. Open the door, then give the machines a reason to quote you.

Want the full Topic Modeler stack?

Five modules + an Enterprise bundle. Foundation projects + ongoing content tooling for AI-search visibility.