About Case Studies
Industries Blog FAQ Contact Get your Free AI Visibility Report →

AI crawlers and robots.txt, a decision with a trade-off.

Blocking AI crawlers is a real option with a real cost. The decision deserves more thought than it usually gets in either direction.

There are two confident camps on this and both are oversimplifying. Blocking AI crawlers protects content from being used without compensation. It also removes you from the answers those systems generate. Which matters more depends on your business model, and it is worth deciding deliberately rather than by default.

What the crawlers do differently

It helps to separate two functions that often use different user agents. Some crawling gathers training data. Other crawling happens live, at the moment a user asks a question, to retrieve current information.

That distinction matters because the trade-off differs. Blocking training crawlers protects your content from being absorbed into a model. Blocking retrieval crawlers removes you from answers being generated right now, which is a much more immediate commercial cost.

Publishers frequently want to block the first and allow the second. That is a coherent position, and whether it is achievable depends on the specific agents each operator documents.

Who should consider blocking

For almost everyone else — service businesses, local companies, ecommerce, B2B — blocking is usually the wrong call. Your content is marketing, and being summarised with attribution is the outcome you want.

Want this audited on your site? Our free AI Visibility Report includes these checks. Request one.

Checking what actually reaches you

Before deciding anything, look at your server logs. Filter for the documented user agents of the major operators and see what is genuinely crawling, how often, and what it fetches.

This frequently produces surprises in both directions: crawlers nobody expected, and absent crawlers explaining why a site is invisible in a particular assistant. It is also the cheapest diagnostic in this whole discipline, and it requires no tool.

Check too whether anything is blocking them unintentionally — a CDN rule, a bot-protection service, or a robots.txt line inherited from a previous team. Unintentional blocking is a more common cause of AI invisibility than most people assume.

Frequently asked questions

Should I block AI crawlers in robots.txt?

For most service businesses, no. Your content is marketing and being cited is the goal. Blocking makes sense mainly for publishers whose product is the content itself, sites with proprietary research, or anyone under contractual restrictions on the material.

Does blocking AI crawlers hurt SEO?

It does not directly affect traditional search rankings, since those crawlers are separate. It does remove you from AI-generated answers, which is an increasingly significant discovery channel in its own right.

How do I know which AI crawlers visit my site?

Check server logs for the documented user agents of the major operators. This is the cheapest diagnostic available and frequently reveals both unexpected crawlers and unintentional blocking by a CDN or bot protection service.

Call now