The technical floor, below which nothing else works.

Most AI visibility problems are content problems. A meaningful minority are a rendering setting nobody checked.

The technical floor, below which nothing else works.

Before any content or entity work matters, a system has to be able to reach your pages and read them. That sounds trivial and is a genuine cause of invisibility often enough to be worth checking first, because it is cheap to check and expensive to assume.

The floor

Reachability

Robots.txt allows the crawlers you want. No CDN or bot-protection rule silently blocking them. No geo-restrictions on regions the crawlers operate from.

Rendering

Content present in the HTML rather than assembled client-side. Some AI crawlers render JavaScript inconsistently or not at all, so content that appears only after execution may simply not exist to them.

Speed and stability

Slow or intermittently failing pages get fetched less reliably. A page timing out during a crawl is indistinguishable from a page that does not exist.

Stable URLs

Content moving between URLs without redirects breaks accumulated understanding, and repeated moves make a site harder to model.

AI CRAWLERS BY PURPOSECompanyModel trainingSearch indexUser-triggered fetchOpenAIGPTBotOAI-SearchBotChatGPT-UserAnthropicClaudeBotClaude-SearchBotClaude-UserPerplexity(none listed)PerplexityBotPerplexity-UserGoogleGoogle-Extended*Googlebotuser fetchersMicrosoft(none listed)Bingbot(Copilot via Bing)Blocking a training bot does not remove you from search answers; blocking a search bot does.*Google-Extended is a robots.txt token, not a separate crawler. Check each vendor's docs for current names.
Major AI crawlers grouped by what they do. Names change; confirm against each vendor's documentation before editing robots.txt.

The checks worth running

Want a baseline before deciding anything?

Our free AI Visibility Report gives you one, with no obligation.

Request one →

The crawlers to look for in your logs

Each major operator documents its user agents, and most separate training, search indexing and user-triggered fetching. OpenAI lists GPTBot, OAI-SearchBot and ChatGPT-User. Anthropic lists ClaudeBot, Claude-SearchBot and Claude-User. Perplexity lists PerplexityBot and Perplexity-User. Google uses Googlebot for search (including AI Overviews) and a separate Google-Extended robots.txt token to control use of content for Gemini models.

Search your access logs for each name and look at the response codes. A run of 403s or 429s against one operator's search crawler is the signature of a firewall rule, and it explains a lot of “why are we never cited in X” questions.

Why this gets missed

These are infrastructure settings, usually owned by a different team from the one worrying about visibility. A bot-protection rule added for security reasons can remove a site from AI answers entirely, and nobody connects the two events because they are months apart and in different systems.

It is worth asking the infrastructure team directly rather than inferring from the outside. Ten minutes of conversation frequently explains months of unexplained invisibility.

Rendering, tested properly

The quickest rendering test is to view the raw HTML of a key page (not the inspector, which shows the rendered DOM) and search for a distinctive sentence from the main content. If it is not there, the content is assembled by JavaScript.

Google renders JavaScript reliably, so this may never have hurt your rankings. Many AI fetchers are simpler and read the HTML they receive. Server-side rendering or static generation for important pages removes the question entirely, and it usually makes pages faster for people too.

Faster discovery

Being crawlable is the floor; being discovered quickly is the next step up. Keep an accurate XML sitemap with real lastmod dates, submit it in Google Search Console and Bing Webmaster Tools, and implement IndexNow so Bing and other participating engines hear about changes as they happen.

Because Copilot, and several other assistants, draw on Bing's index, Bing Webmaster Tools deserves more attention than most teams give it. It also now reports how often your pages are cited in AI answers, which is a useful check that the technical floor is holding.

Frequently asked questions

What technical requirements matter for AI search?

Crawler access unblocked by robots.txt, CDN or WAF rules; content present in HTML rather than assembled client-side; adequate speed and stability; stable URLs; and valid structured data.

Does JavaScript rendering affect AI search visibility?

It can. Some AI crawlers render JavaScript inconsistently or not at all, so content that only appears after execution may be invisible to them even though it is fine for traditional search.

Why might a site be invisible in AI search despite good content?

Frequently an infrastructure setting: a bot-protection or CDN rule blocking AI crawlers, added for security reasons by a different team months earlier. It is worth asking directly rather than inferring.

Does Google-Extended affect AI Overviews?

No. Google-Extended controls whether content is used for Gemini model training and grounding in some Google AI products; AI Overviews are part of Search and use Googlebot.

Do AI crawlers respect robots.txt?

The major operators say their crawlers do. User-triggered fetchers can behave differently because a person requested the page, so check each operator's documentation.

Sources & further reading

  1. Introduction to robots.txt — Google Search Central
  2. Overview of OpenAI crawlers — OpenAI Platform docs
  3. Perplexity crawlers — Perplexity docs
  4. Does Anthropic crawl data from the web, and how can site owners block the crawler? — Claude Help Center
  5. Google's common crawlers (incl. Google-Extended) — Google for Developers
  6. Bing Webmaster Guidelines — Microsoft Bing
  7. IndexNow protocol — IndexNow.org
Call now