Is Your Website Blocking AI Crawlers? How to Check in 90 Seconds
Short answer: quite possibly, and nobody told you. AI tools like ChatGPT, Claude and Perplexity read the web using their own named crawlers. A small file on your website called robots.txt tells each of those crawlers yes or no. Over the last two years, hosting companies, security plugins and Cloudflare have all added AI crawler blocking, and plenty of them switched it on by default. If yours is blocking, you cannot be recommended by those tools. Not ranked lower. Not there at all.
Here is how to check, and how to fix it if the answer is bad.
What an AI crawler is
An AI crawler is a bot that reads web pages so an AI system can use what it finds. Each one has a name.
These sit alongside Googlebot, which you already know about. Blocking Googlebot would be an obvious disaster. Blocking GPTBot is a quiet one, because nothing visibly breaks.
| Crawler | Belongs to | What it feeds |
|---|---|---|
| GPTBot | OpenAI | ChatGPT training and search |
| OAI-SearchBot | OpenAI | ChatGPT search results |
| ClaudeBot | Anthropic | Claude |
| PerplexityBot | Perplexity | Perplexity answers |
| Google-Extended | Gemini and AI Overviews training | |
| Applebot-Extended | Apple | Apple Intelligence |
| CCBot | Common Crawl | Feeds many downstream AI datasets |
| Bingbot | Microsoft | Bing index, which Copilot relies on |
Why this comes before everything else
Every other thing you might do to improve your AI visibility assumes the machines can read your website. Schema markup, entity statements, FAQs, business listings. All of it is you handing over information.
If the crawler is blocked at the door, none of that information gets picked up. You can spend a weekend on the rest of it and change nothing.
This is why crawler access is the first check, and why it is worth doing before you read anything else about AI search.
How to check your robots.txt file
Open a browser and type your domain followed by /robots.txt. So yourbusiness.co.nz/robots.txt.
Read what comes up. It will be plain text, usually short.
Look for the word Disallow underneath any of the crawler names in the table above.
A line that says User-agent: GPTBot followed by Disallow: / means ChatGPT cannot read a single page of your site.
A line that says Disallow: with nothing after it means the opposite. That one allows everything. The difference is one character, which is why people misread it.
If your robots.txt file does not mention any AI crawlers at all, that is good news. No rule means allowed.
Check your security layer too
Robots.txt is only half of it. Robots.txt is a request. A firewall is a wall.
Cloudflare. Log in, go to Security, then Bots. Look for a setting mentioning AI Scrapers and Crawlers or Block AI Bots. Cloudflare turned this on by default for new sites in 2024, so a lot of owners have it without ever choosing it.
WordPress. Check whichever security plugin you installed and forgot about. Wordfence, All In One Security and similar have added AI bot blocking options.
Shopify, Wix, Squarespace. Look in settings for a robots.txt editor. Several platforms have added AI blocking toggles in the last year.
Your host. Some managed hosts apply bot rules at server level that never appear in your robots.txt. If the file looks clean but you are still not being read, ask your host directly whether AI crawlers are blocked.
How to fix it
Delete any Disallow rule sitting under GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended, Applebot-Extended or CCBot.
Switch off AI scraper blocking in Cloudflare or your host security settings.
Save, reload your robots.txt in a browser, and read it again to confirm the change actually took.
Give it a few weeks. Crawlers revisit on their own schedule and nothing happens instantly.
A prompt to read the file for you
Paste your robots.txt into any AI tool with this:
Below is the robots.txt file from my website. Tell me whether each of these is allowed or blocked: Googlebot, Bingbot, GPTBot, ClaudeBot, PerplexityBot, Google-Extended, Applebot-Extended, CCBot. Give me a two column list. For anything blocked, quote the exact line I need to delete. Do not explain how robots.txt works.
Thirty seconds, and you get a verdict instead of squinting at syntax.
This is a real decision, not just a fix
Worth saying plainly. Blocking AI crawlers is a legitimate choice. It keeps your content out of AI training, which matters if your content is your product. Publishers, course creators and people selling written work often block on purpose.
Allowing them is what gets you recommended.
You cannot have both. The point of this check is not that everyone should allow everything. The point is that you should be choosing, rather than discovering in two years that a plugin decided for you.
For most local and service businesses, the content is not the product. The business is. Those businesses should almost always allow.
Common mistakes
Assuming no news is good news. Nothing breaks visibly when a crawler is blocked. There is no error, no email, no drop in a report you already look at.
Fixing robots.txt but leaving Cloudflare on. The firewall wins. Check both.
Blocking Bingbot by accident. Bing feeds Copilot and parts of ChatGPT search. People who wrote Bing off a decade ago sometimes deprioritise it in their bot settings without realising what it now feeds.
Expecting instant results. Crawl frequency for a small site can be weeks. Fix it and move on to the next thing.
Frequently asked questions
Does blocking AI crawlers affect my Google rankings? Blocking Google-Extended does not affect normal Google Search rankings. It only affects Gemini and AI training use. Blocking Googlebot absolutely does affect rankings, and that is a different setting.
Will allowing AI crawlers mean my content gets stolen? Your content may be used in training and may be summarised in answers. That is the trade. In exchange, you become eligible to be named when someone asks for a business like yours. Decide which matters more for your business model.
I do not have a robots.txt file at all. Is that bad? No. No file means no restrictions, so everything is allowed. You may want one eventually for other reasons, but for this check you pass.
How often should I recheck this? Every six months, and immediately after any host migration, security plugin install, or Cloudflare setup. Those three events are what change it without warning.
Can I allow some AI crawlers and block others? Yes. Each one is named separately in robots.txt, so you can allow OpenAI and block Common Crawl, or any combination you want.
This is check one of seven
Crawler access is the first of seven things that decide whether AI tools can find you, trust you and name you. The other six are your business details, your schema markup, Search Console, your Google, Bing and Apple listings, your profile consistency, and your FAQ.
Get Named By AI is the full 7-Point AI Visibility Check. Seven checks in the order they have to be done, a scorecard so you know where you stand, and nine copy and paste prompts that do the work for you. 21 pages, $17, and about an afternoon of your time. Get Named By AI for $17