Skip to content
CactusLaunch

How ChatGPT, Perplexity, and Claude Find Your Website (and the robots.txt Mistakes That Hide It)

Every AI assistant that cites websites reaches them through a crawler, and each company documents exactly which bot does what. Here is the decision table built only from those official pages, the difference between training bots and search bots, the CDN and firewall settings that silently block all of them, and the honest verdict on llms.txt.

Brian Simmons · Founder, CactusLaunch 6 min read
A row of clean stone stepping stones crossing a shallow, still desert spring toward a lit doorway at first light

When an AI assistant recommends a business, it found that business through a crawler, and the company running the assistant tells you which crawler and what it does. This is one of the few areas of AI search where there is no mystery: OpenAI, Perplexity, Anthropic, Google, and Microsoft all publish bot documentation. Almost nobody reads it, which is how businesses end up blocking the bot that would have cited them while allowing the one they meant to stop.

Here is the decision table built only from those official pages, the settings that silently block everything, and the honest verdict on the file everyone keeps asking about.

How a business ends up cited in an AI answer

  1. Your website Clear entity, structured data, real content
  2. Crawlers fetch it Googlebot, OAI-SearchBot, PerplexityBot
  3. Indexes and models Facts are extracted and stored
  4. Someone asks a question “Who does X near Y?”
  5. Answer is assembled From sources the system trusts
  6. Citation or omission Legible sites get named; vague ones don’t
Nobody controls the last step directly. The first step is the one a business owns completely.

The crawlers, by what they actually do

Every operator separates at least two jobs: crawling for training (building the model) and crawling for search (surfacing and citing live web pages in answers). Several add a third: fetching a page because a user asked for it.

OperatorBotJobHonors robots.txt?What blocking it does
OpenAIOAI-SearchBotSurfaces websites in ChatGPT searchYes (changes take ~24 hours)Per OpenAI, sites opted out will not be shown in ChatGPT search answers
OpenAIGPTBotTrainingYesKeeps your content out of model training; no effect on search visibility
OpenAIChatGPT-UserFetches a page a user asked aboutMay not applyNot used to decide search inclusion
PerplexityPerplexityBotSurfaces and links websites in Perplexity answersYes; Perplexity recommends allowing itRemoves you from Perplexity answers
PerplexityPerplexity-UserFetches a page a user asked aboutGenerally ignores robots.txt—
AnthropicClaude-SearchBotSearch quality for ClaudeYesReduces visibility in Claude’s search
AnthropicClaude-UserFetches a page a user asked aboutYesAnthropic says blocking may reduce visibility for user-directed search
AnthropicClaudeBotTrainingYes (and Crawl-delay)Keeps content out of training
GoogleGooglebotSearch, including AI Overviews and AI ModeYesRemoves you from Google Search entirely
GoogleGoogle-ExtendedControls training use in other Google productsYesDoes not affect Search or its AI features
MicrosoftBingbotBing search, which feeds Copilot and partnersYesRemoves you from Bing and Copilot

The pattern: allow the search bots and the user fetchers; decide separately about the training bots. A business that wants to be found has no reason to block OAI-SearchBot, PerplexityBot, Claude-SearchBot, or Bingbot. Whether to block GPTBot, ClaudeBot, or Google-Extended is a content-rights decision with no visibility cost either way.

A robots.txt that reflects that looks like this:

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: *
Allow: /
Sitemap: https://example.com/sitemap-index.xml

Or, if you’re happy for your content to be used in training, simply allow everything. What you should never do is a blanket Disallow: / under a wildcard, or a block on a search bot because it “sounds like AI.”

The blocks that aren’t in robots.txt

Here is the finding from our own audits that matters most: the sites most often invisible to AI systems have a perfectly reasonable robots.txt. They’re blocked somewhere else.

CDN bot protection. Cloudflare, and other CDNs, offer bot-fighting modes and AI-crawler controls that can challenge or block automated traffic wholesale. Turned on with default settings, some configurations block the search bots above along with the scrapers. Check your CDN’s bot settings and, where the CDN offers per-bot controls, allow the search and user-fetch bots explicitly.

Web application firewalls. Rules that block “unknown” user agents or rate-limit aggressively can return errors to crawlers. The crawler doesn’t complain; it just records that your site doesn’t respond.

Security plugins. WordPress security plugins with “block bad bots” features maintain their own lists, which are sometimes out of date and sometimes overbroad.

Hosting-level blocks. Some managed hosts block crawlers by default on staging or on lower tiers.

JavaScript-only content. Not a block, but the same result: if your service details render only after a script runs, some crawlers see an empty page. Content should be in the HTML.

The test is simple. Fetch a key page with each bot’s user-agent string (any HTTP tool can do this) and confirm you get the page, with your content in it, not a challenge page or an error. Do it after any CDN, firewall, or plugin change.

The llms.txt question, settled as far as it can be

llms.txt is a proposed convention — a markdown file at your site’s root summarizing who you are and where your key pages live, specified at llmstxt.org since 2024 and revised in 2026. The idea is reasonable. The evidence for its effect is not there.

  • Google says its Search doesn’t use these files and that they neither help nor harm; a Google search advocate stated publicly in 2025 that no AI system used llms.txt.
  • No major AI company — OpenAI, Anthropic, Perplexity, Microsoft — has documented reading llms.txt for search or answers. Several publish llms.txt files for their own developer documentation, which is publishing, not consuming.
  • Ahrefs measured it in June 2026 across more than 137,000 domains: about 28% had the file, 97% of those files received zero requests, and AI retrieval bots accounted for around 1% of all requests. Their verdict was that it is largely decoration.

Our position: it’s cheap, it’s harmless, and it has a legitimate niche in developer documentation for coding tools. This site keeps one because it costs nothing and states our facts accurately. It is not a lever, we don’t sell it as one, and if a proposal for your site leads with llms.txt, ask what else is in it.

What actually makes an AI answer cite you

Strip away the files and the acronyms and what remains is short:

  1. Reachable. Search bots allowed at every layer — robots.txt, CDN, firewall, plugins.
  2. Readable. Content in HTML, pages that load fast and completely.
  3. Legible. A consistent, plain description of who you are, what you do, and where — on your site, your Business Profile, and the directories. Schema markup states the same facts in machine form; it helps understanding, not placement.
  4. Worth citing. Pages that answer specific questions specifically. An answer engine quotes the page that says what a service involves and costs, not the page with the slogan.
  5. Indexed by Google and Bing. Most AI search systems lean on or resemble those indexes. Regular SEO is the foundation; Google AI Overviews and AI Mode covers Google’s side specifically.

None of this guarantees a citation. Nobody publishes a selection formula, and answers change between sessions. What it guarantees is that you’re a candidate, which is the part a business controls.

Checking your own site in ten minutes

  1. Open yourdomain.com/robots.txt and read it against the table above.
  2. Check your CDN’s bot-protection settings for anything blocking search bots.
  3. Fetch a service page as OAI-SearchBot and as PerplexityBot and confirm the content comes back.
  4. Ask ChatGPT, Perplexity, and Google’s AI Mode what your business does and who does your service in your area. Note what’s wrong.
  5. Fix the facts at the source — your site and your profiles — rather than hoping the answer improves on its own.

That audit, done properly and followed by the content and consistency work it reveals, is what our AI search optimization service consists of. If you’d like it done for you, that’s where to start; if you’d like the wider picture first, How AI Search Changes Website SEO has it.

Questions people ask

Should I block AI crawlers to protect my content?

Decide bot by bot, because they do different things. Blocking a training bot (GPTBot, ClaudeBot, Google-Extended) keeps your content out of model training and costs you nothing in visibility. Blocking a search bot (OAI-SearchBot, PerplexityBot, Claude-SearchBot) removes you from that assistant's answers. For a business that wants customers to find it, the search bots should be allowed.

Does robots.txt actually stop them?

The search and training bots from OpenAI, Perplexity, Anthropic, and Google all document that they honor robots.txt. The user-request fetchers — ChatGPT-User, Perplexity-User — fetch a page because a person asked for it, and both companies say those may not follow robots.txt rules. Robots.txt controls the crawling; it doesn't control what a user chooses to read.

Do I need to submit my site to ChatGPT or Perplexity?

No. There is no submission process; the crawlers find sites through links, sitemaps, and existing indexes, the same way Googlebot does. What you can do is make sure nothing is blocking them and that your site is worth citing.

Does llms.txt help?

Not in any way anyone has demonstrated. Google says its Search doesn't use it. No major AI company has documented reading it for search or answers. Ahrefs' 2026 measurement across more than a hundred thousand domains found the files almost never requested. It's harmless and cheap, and we keep one on this site for exactly that reason — but it is not a lever, and anyone selling it as one is mistaken.

Why does ChatGPT describe my business wrong?

Because it assembled its answer from whatever it could find — old directory listings, an outdated page, a competitor's site — and your own site didn't state the facts clearly enough to win. The fix is on your side: an accurate, consistent description of who you are, what you do, and where, on your site and your profiles, reachable by the search bots.

This article is part of the ai search visibility library. If it describes a problem you have, the service behind it is here: see ai search optimization.

Make sure AI describes your business accurately.

Entity clarity, crawl access, and answer-shaped content — the groundwork, honestly framed.

Prefer to talk it through first? Use the chat button — a real person replies.