AI Crawler Access & llms.txt Test
The AI Crawler Access & llms.txt Test checks two things that decide whether AI assistants can find and cite your website.
- Results in seconds
- Pass / fail + fix guidance
- No account required
The AI Crawler Access & llms.txt Test checks two things that decide whether AI assistants can find and cite your website. First, it reads your robots.txt and separates AI search crawlers (OAI-SearchBot for ChatGPT Search, PerplexityBot) from training-only crawlers (GPTBot, ClaudeBot, Google-Extended, CCBot and others). The test fails only when search crawlers are blocked. Cloudflare’s default “Disallow AI Training” is treated as informational, not a fail. Second, it fetches /llms.txt, the emerging standard file that gives models a curated summary of your site.
What This Tool Checks
- /llms.txt exists at the root of your domain and is readable
- robots.txt rules for AI search crawlers (OAI-SearchBot, PerplexityBot)
- robots.txt rules for training-only crawlers (GPTBot, Google-Extended, ClaudeBot, CCBot)
- Cloudflare managed robots.txt / Content-signal ai-train=no (reported, not failed)
- Blanket Disallow rules that unintentionally block all crawlers, including search
Why It Matters for SEO
A meaningful share of product and information searches now start inside ChatGPT, Perplexity, Claude and Google AI Overviews rather than a classic search results page. Those systems can only cite content their search crawlers can read. If robots.txt blocks OAI-SearchBot or PerplexityBot, your brand does not appear in those answers. Blocking GPTBot or Google-Extended only opts out of model training — the current Cloudflare recommended default — and does not hide you from AI search. llms.txt does not gate access; it tells models which pages best represent you.
How to Fix It
Open your robots.txt and look for groups naming OAI-SearchBot or PerplexityBot. Remove Disallow: / for those search crawlers if you want to be cited in ChatGPT Search and Perplexity. You can leave GPTBot, Google-Extended, ClaudeBot and similar training crawlers disallowed if that is your licensing position — Cloudflare now recommends that for most sites. Then create an llms.txt file: a short markdown document at https://yourdomain.com/llms.txt with an H1 for your site name, a one-line blockquote summary, and a list of your most important pages with one-line descriptions.
How It Works
We fetch your robots.txt and evaluate it against each known AI user-agent, including wildcard and default (*) groups. In parallel we request /llms.txt at your domain root. The test passes when AI search crawlers can access your site and an llms.txt file is present; it warns when llms.txt is missing; and it fails only when search crawlers are blocked. Training-only Disallow rules do not fail the test.
Common Mistakes to Avoid
- A blanket "User-agent: * Disallow: /" left over from staging that blocks every crawler, including AI search
- Treating Cloudflare’s “Disallow AI Training” as an accidental AI block — it is the new default
- Blocking OAI-SearchBot or PerplexityBot while thinking you only blocked training
- Publishing llms.txt with broken links or pages that themselves block crawlers
- Assuming llms.txt controls access — it is a guide, robots.txt is the gate
Quick Checklist
- robots.txt does not block OAI-SearchBot or PerplexityBot
- /llms.txt exists and returns 200 with markdown content
- llms.txt lists your most important, canonical pages
- Training-only crawlers (GPTBot, Google-Extended) blocked or allowed on purpose
- Re-checked after every robots.txt or Cloudflare bot-settings change
Put your whole site on autopilot — SEO and AI search
PositionMySite monitors every signal on this page across your entire website 24/7 — plus keyword rankings, competitor moves and AI-search readiness (llms.txt, schema, ChatGPT & Gemini visibility). When something breaks, you know before Google does.
Every feature unlocked · No commitment · Cancel anytimeFrequently Asked Questions
A proposed standard (llmstxt.org) for a markdown file at /llms.txt that gives large language models a concise, curated overview of your site — what it is, and which pages matter most. Think of it as a sitemap written for AI instead of crawlers.
Blocking GPTBot opts your content out of OpenAI model training. ChatGPT Search uses OAI-SearchBot; live browsing uses ChatGPT-User. Those are the tokens that decide whether you can still be cited in ChatGPT answers.
Most sites should allow search crawlers (OAI-SearchBot, PerplexityBot) and can disallow training-only crawlers (GPTBot, Google-Extended, CCBot). That is Cloudflare’s recommended split as of September 2026.
No — AI systems can cite any page their search crawlers reach. llms.txt improves the quality of citations by pointing models at your best, canonical pages, and signals that your site is AI-friendly.
The robots.txt token Google uses for Gemini training access. Blocking it does not affect normal Google Search rankings. It is a training opt-out, not a search block.
Cloudflare replaced a single “Block AI Bots” switch with Search, Training, and Agent controls. The recommended Training setting is “Disallow AI Training,” which writes robots.txt rules for GPTBot, Google-Extended and similar training crawlers while leaving Googlebot and AI search crawlers allowed.