Skip to content
EXVOQ Be Found. Be Chosen.

blog · AI Visibility

Is Your Website Blocking AI Crawlers? How to Check robots.txt

By Amit · 29 July 2026 · 6 min read

If your robots.txt file blocks GPTBot, ClaudeBot or PerplexityBot, AI assistants cannot read your pages, so they will never quote your business in an answer. You can check this yourself in two minutes: open yoursite.com/robots.txt in any browser and look for Disallow lines sitting under those crawler names. If you find them, the fix is to allow GPTBot in robots.txt along with the other AI agents, which takes about ten minutes and costs nothing. Many owners never look at this file, and plenty of hosting platforms and security plugins add the blocks without telling anyone.

What is robots.txt and why does it decide whether AI can see you?

robots.txt is a plain text file that sits at the root of your website. Every automated visitor, from Googlebot to ChatGPT's crawler, reads it first and treats it as a set of house rules: which parts of the site it may fetch, and which it should leave alone. It is a request, not a lock, but the crawlers run by the major AI companies do follow it.

That makes it the first gate in the whole chain. Good content, clean design and correct schema markup do nothing if the crawler was turned away at the door. This is also why a site can rank respectably on Google and still be completely absent from AI answers. Google has been crawling you for years under a different agent name. The AI agents are newer, and they are the ones getting blocked.

Which AI crawlers should I be looking for?

You only need to recognise a handful of names. When you open your robots.txt, scan for these:

  • GPTBot — OpenAI's crawler, the one that builds the index behind ChatGPT.
  • OAI-SearchBot — OpenAI's separate agent for search results inside ChatGPT.
  • ChatGPT-User — fetches a page live when a user asks ChatGPT to look something up.
  • ClaudeBot — Anthropic's crawler for Claude.
  • PerplexityBot — Perplexity's crawler.
  • Google-Extended — not a crawler at all, but a control token. Googlebot still crawls your site normally; this token tells Google whether that content may be used in Gemini and related AI products. Blocking it does not affect your normal Google ranking, but it does take you out of some AI answers.

If any of these names appears with a Disallow rule pointing at your whole site, that assistant is reading nothing from you.

How do I check my own robots.txt in two minutes?

Type your domain into a browser and add /robots.txt to the end. So exvoq.com becomes exvoq.com/robots.txt. You will see one of three things.

  • A page not found error. No robots.txt exists. Crawlers treat this as permission to read everything, so nothing is blocked. Not ideal for other reasons, but not an emergency.
  • A short file with no AI crawler names in it. Usually fine. If the only rules are about admin folders or a checkout page, the AI agents are free to read your real content.
  • A file that names GPTBot, ClaudeBot, PerplexityBot or Google-Extended with a Disallow line. This is the problem case, and it is more common than owners expect.

One detail trips people up. A Disallow line followed by nothing at all means allow everything. A Disallow line followed by a single forward slash means block everything. Those two lines look almost identical and mean opposite things, so read them slowly.

How do I allow GPTBot in robots.txt without opening up things I want private?

The goal is not to remove your robots.txt. It is to make sure the AI agents get the same access Google already has. In practice that means deleting the blanket Disallow rules aimed at the AI crawler names, while keeping any rules that protect genuinely private areas such as admin logins, cart pages or internal search results.

Where you edit the file depends on your platform. On WordPress it is usually managed by an SEO plugin such as Yoast or Rank Math, under a tools or file editor section. On Shopify and Wix there is a robots.txt setting in the site or SEO preferences. On a custom or static site it is a real file in your public folder, and your developer can change it in a minute. On some setups a security layer or CDN adds its own AI blocking rules on top, so if you edit the file and the blocks reappear, check there next.

After you save, reload yoursite.com/robots.txt and confirm the change actually shows. Caching sometimes serves the old version for a while.

Why would my site be blocking AI crawlers if I never set that up?

Almost nobody does this on purpose. It usually arrives through one of these routes:

  • A hosting platform or CDN added AI bot blocking as a default protective feature, sometimes framed as saving bandwidth or protecting content.
  • A security plugin bundles AI agents into its general bot blocking list.
  • A developer copied a robots.txt template from a forum post written when blocking AI training crawlers was the popular advice.
  • The block was added deliberately at some point by someone who has since left, and nobody revisited the decision.

The reason matters less than the result. If you have wondered why a competitor keeps getting named by ChatGPT and you do not, this is worth ruling out before anything else. Our post on how to check if your business appears in ChatGPT walks through testing the visibility side once access is sorted.

Does allowing AI crawlers mean giving away my content for free?

This is a fair question and the honest answer is that it depends on your business. If you sell the content itself, such as a paid course or subscription publication, blocking training crawlers is a defensible choice. Most small businesses are in the opposite position. Their website is marketing material whose entire job is to be found and repeated. Being quoted by an AI assistant, with your name attached, is free distribution to someone who is actively asking for what you sell.

There is also a middle path. You can allow the search and live-fetch agents that produce cited answers while blocking the pure training crawlers. It is more fiddly to maintain, and for most MSMEs the simpler open approach is the better trade. The wider structural picture is covered in our note on how AI is changing website design.

FAQ

How long after fixing robots.txt will AI assistants mention my business?

There is no fixed timeline, because each crawler returns on its own schedule and every assistant refreshes its index differently. Treat the fix as removing a blocker rather than as a switch that produces results on a date.

Do I need a separate rule for every AI crawler name?

No. If your file has no rules targeting those agents at all, they are already allowed. You only need to name them if you are removing an existing block or deliberately treating one differently from the rest.

Can blocking AI crawlers hurt my normal Google ranking?

Blocking Google-Extended does not affect standard Google Search ranking, since Googlebot handles that separately. But a badly written robots.txt can block Googlebot by accident, which does hurt, so it is worth reading the whole file rather than only the AI lines.

EXVOQ's free instant audit checks your robots.txt for AI crawler access, along with your structured data and content format, in about 15 seconds. Try it on your own site.

free · 48-hour turnaround

Find out what AI says when customers ask about your market.

Your AI Visibility Audit shows which assistants know your business, what they say, and exactly what to fix. No credit card, no pitch.

Get your free AI visibility audit