Table of Contents
- What Are AI Crawlers and Which Ones Actually Matter?
- Why Are AI Crawlers Blocked When Your Content Is Fine?
- How Do You Check Whether AI Crawlers Can Reach Your Site?
- How Do You Unblock AI Crawlers Without Giving Away Everything?
- What Should You Do Once AI Crawlers Can Reach Your Pages?
- Open the Door First, Then Give the Bots Something Worth Citing
- Frequently Asked Questions
- Your content can be invisible in AI search because a firewall or CDN default blocks AI crawlers, even when the pages themselves are great.
- Checking crawler access is the mandatory first step of technical SEO for AI search, because every other optimization depends on it.
- Allow the search and user-fetch bots you want, then verify at the server level instead of trusting robots.txt alone.
If your pages never show up in AI answers, the first thing to check is whether AI crawlers can actually reach them. Website security systems and crawl rules block these bots more often than most site owners realize, and a page a bot cannot fetch is a page AI search cannot cite. The short version: verify access at two layers (your robots.txt file and your firewall or CDN), allow the bots that send you visibility, and test again.
This applies to every site, but it hits hardest when you run startup digital marketing on a lean budget and lean on your host’s or CDN’s default settings. Those defaults were written for a web that existed before AI search, and some of them quietly close the door. Below you will find which bots matter, why blocks happen, how to check your own site, and how to fix what you find.
What Are AI Crawlers and Which Ones Actually Matter?
AI crawlers are automated bots that fetch web pages on behalf of AI products. They do three different jobs: collecting training data, building the search indexes behind cited answers, and fetching a page live when a user asks about it. Only the second and third jobs decide whether your content appears in an AI answer.
The big vendors split those jobs across separate bots, which is good news because you can control each one on its own:
- OpenAI: GPTBot gathers content that may be used for model training, OAI-SearchBot builds the index behind ChatGPT search, and ChatGPT-User fetches a page when a person asks ChatGPT to look at it.
- Anthropic: ClaudeBot is the training crawler, Claude-SearchBot supports search, and Claude-User handles live fetches on a person’s behalf.
- Perplexity: PerplexityBot indexes sites for Perplexity’s search engine.
- Google: Googlebot indexes pages for Search and its AI features, while Google-Extended is a robots.txt token that controls whether your content can be used to train Gemini models.
OpenAI’s own crawler documentation states that each of these settings is independent. It also says that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though they can still appear as navigational links.
That detail trips people up. Blocking GPTBot to keep your writing out of training data does not remove you from ChatGPT search. Blocking OAI-SearchBot does.
Why Are AI Crawlers Blocked When Your Content Is Fine?
Most blocks come from infrastructure, not content. Firewalls, CDNs, bot protection, and rate limiters decide whether a request ever reaches your server, and they make that call before your robots.txt file or your writing gets a vote.
That is why the most common cause of AI invisibility is an infrastructure default rather than a content problem. Here is how it usually happens:
- Robots.txt is a request, not a lock. It states your rules, but your CDN or firewall enforces its own. One developer who built an AI bot access checker described sites whose robots.txt welcomed GPTBot while the bot received a 403 Forbidden response from the infrastructure in front of the site.
- Bot protection is switched on by default. Security tools tuned to stop scrapers often treat all AI bots alike, including the search bots you want.
- Rate limits are too tight. A crawler fetching many pages quickly can trip limits designed for abusive traffic.
- Your CDN changed its defaults. According to published reports of Cloudflare’s July 1, 2026 announcement, its new defaults took effect on September 15, 2026. Training and agent crawlers are blocked by default on ad-supported pages for new domains, new sites on existing accounts, and free-tier accounts that had not changed their settings, while search crawlers stay allowed. The same reports say mixed-purpose crawlers, Googlebot included, follow the most restrictive rule that applies. If you use Cloudflare, open your dashboard and confirm instead of assuming.
The pattern repeats across providers: a setting nobody consciously chose decides what AI search can see. So the diagnosis starts with access, not with rewriting your content.
How Do You Check Whether AI Crawlers Can Reach Your Site?
Check access in layers, from the cheapest test to the most reliable one. Together these five steps tell you what your rules say and what your server actually does.
- Read your live robots.txt. Open yourdomain.com/robots.txt and look for a blanket Disallow: /, named rules for the bots above, and any block your CDN or host injected without you writing it.
- Review your CDN and firewall settings. Look for AI bot toggles, bot management modes, WAF rules, and rate limits. This is where the silent defaults live.
- Test with curl. Request a page as a plain visitor, then again with an AI bot’s user agent string, and compare the status lines:
curl -I https://yourdomain.com/
curl -I -A “Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-SearchBot/1.0; +https://openai.com/searchbot” https://yourdomain.com/
A 200 on both is encouraging. A 403 or 429 on the bot request points to a rule keyed on user agent. One caveat: real bots arrive from published IP ranges, and some firewalls check both, so a clean curl result is a good sign rather than proof.
4. Check your server logs. Search for the bot names and count the status codes they received (this assumes a standard combined log format):
grep -E “GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|PerplexityBot” access.log | awk ‘{print $9}’ | sort | uniq -c
Plenty of 200s means the bots are getting through. A pile of 403s or 429s means something is stopping them. OpenAI publishes the IP ranges for its bots, so you can confirm that a request claiming to be OAI-SearchBot really came from OpenAI.
5. Run an audit. The SEMrush SEO Toolkit‘s Site Audit includes a Blocked from AI Search check that shows how many audited pages your robots.txt blocks from AI search bots, including GPTBot. It replaces the first manual step with a single view across hundreds of pages. Because it reads robots.txt rules, you still need steps 2 to 4 to catch firewall-level blocks.
Here is why this check deserves to go first. Structured data, answer-first formatting, and fresh content all assume the bot can fetch the page. If it gets a 403, none of that work is ever seen.
Yet the usual way to verify access today is a pile of curl commands and log-file forensics, so many teams skip it and chase content tweaks instead. Treat crawler access verification as step zero of AI-era technical SEO, the gate everything else depends on.
How Do You Unblock AI Crawlers Without Giving Away Everything?
Decide by purpose, not by vendor. Allow the search and user-fetch bots that can send you visibility, then make a separate, deliberate choice about training bots based on your own content policy.
Here is a robots.txt example that welcomes the search and fetch bots while opting out of OpenAI training. Delete the GPTBot block if you are comfortable with training use.
User-agent: OAI-SearchBot
Allow: /User-agent: ChatGPT-User
Allow: /User-agent: PerplexityBot
Allow: /User-agent: GPTBot
Disallow: /
Robots.txt alone will not fix an infrastructure block, so work through these steps too:
- Allowlist verified IP ranges. OpenAI recommends allowing requests from its published IP ranges in addition to permitting OAI-SearchBot in robots.txt. Allowing by user agent alone is weaker, since anyone can fake a user agent string.
- Adjust bot protection rules. Switch off any blanket AI bot block that catches search crawlers, or add exceptions for the ones you want.
- Loosen rate limits for verified bots. Give known crawlers enough room to fetch your important pages without tripping abuse protection.
- Re-test and watch the logs. Repeat steps 3 and 4 from the checklist, then monitor for a few weeks to see the bots return with 200 responses.
One honest limit: unblocking makes you eligible, and it does not guarantee citations. No setting forces an AI product to quote you.
What Should You Do Once AI Crawlers Can Reach Your Pages?
Once AI crawlers can fetch your pages, your job shifts from access to being worth citing. Four habits do most of the work.
Make the page readable by a bot. Serve your main content in the HTML response, because many AI crawlers do not execute JavaScript the way a browser does. Lead each section with a direct answer, then add detail and an example.
Plan content around real questions. Start from the questions people put to AI tools, then build an SEO content strategy that answers them in clearly labeled sections. Group related posts into clusters so each page supports the others and a reader (or a bot) can move between them easily.
Name your entities consistently. Use the same names for your products, people, and topics on every page, and rely on entity linking between related posts so the relationships are obvious to a machine. Consistency is a cheap trust signal, and it is entirely in your control.
Earn mentions, then protect them. Off-site, well-run PR campaigns create the third-party references that AI systems weigh when deciding which sources to trust. Those references only count if the pages they point to return a clean 200 response to crawlers, so review your existing link building strategies with the same access checks you just learned.
Open the Door First, Then Give the Bots Something Worth Citing
AI search visibility starts with a plain question: can the bots get in? A firewall rule, a CDN default, or a rate limit can answer “no” while your robots.txt says “yes,” and the usual cost is months of invisible content that looks like a quality problem when it is really an access problem.
So run the five checks, allow the bots that send you visibility, and decide on training bots as a separate business choice. Re-check after every change to your CDN, host, or security plugins, because defaults shift without much warning. Then spend your energy on the content, entities, and mentions that make a page worth citing.
Ready to See What AI Crawlers See on Your Site?
Run a Site Audit with the SEMrush SEO Toolkit to find out which of your pages are blocked from AI search bots, then use the steps above to verify your firewall and CDN. It takes the first and most tedious check off your plate.
Frequently Asked Questions
Does robots.txt alone prove AI crawlers can access my site?
No. Robots.txt only states your rules, while a firewall or CDN can still block the request before it arrives.
If I block GPTBot, will I disappear from ChatGPT search?
Not by itself. GPTBot is used for training, while OAI-SearchBot controls whether your pages appear in ChatGPT search.
How often should I re-check AI crawler access?
Check after any change to your CDN, firewall, host, or security plugins. A quarterly check is a sensible baseline.
Disclosure: This post is promotional content and contains affiliate links. If you sign up for or buy a product through a link marked “affiliate link,” mariaisquixotic earns a commission at no extra cost to you. Our opinions and recommendations are our own.
Maria is a digital marketing professional, specializing in content marketing and SEO. She's a neurodivergent who strives to raise awareness, and overcome the stigma that envelopes around mental health.






No Comment! Be the first one.